How restorability is verified, in detail
This page is for operators who want to know precisely what a green rating proves. It describes the implementation of Restow 0.1.0, with the source files listed at the end. The plain-language overview is Recovery readiness.
Two rules hold everywhere:
- A rating belongs to one backup. A check records the snapshot it read. A newer backup is Not verified until a check of that backup has run, whatever an older one scored.
- Green needs a read-back. A backup job that finished is not a pass. Only content that was read back and matched what the backup recorded at the time makes a backup Ready.
Mailboxes, OneDrive and IMAP accounts
Section titled “Mailboxes, OneDrive and IMAP accounts”What a backup records
Section titled “What a backup records”| Record | Algorithm | Of what | Where |
|---|---|---|---|
| Chunk id | HMAC-SHA-256 with a key of the tenant | the plaintext of each content-defined chunk | the snapshot manifest and the chunk index (PostgreSQL) |
| Sealed chunk | AES-256-GCM with the tenant’s data key; the chunk id is bound into the sealed header as GCM associated data | each chunk | inside a pack file on the storage target |
| Pack hash | SHA-256 | the whole pack file (header, chunks, index, footer); the pack footer carries the same hash | the pack catalog (PostgreSQL), together with the pack size |
| Object hash | SHA-256 | the reassembled bytes of each item (a mail, a file, a calendar or contact item) | the manifest entry, together with the size and the ordered chunk ids |
| Manifest | AES-256-GCM, sealed with the tenant key | the list of items of one snapshot | the storage target |
The restore check
Section titled “The restore check”Each check reads the latest completed snapshot of one protected object:
- Manifest. The sealed manifest is loaded and opened. If it comes back but does not decode (damaged, or a key that does not open it), the result is Not restorable (
manifest_unreadable). If the storage does not return it at all, the check is not completed (see Red only with evidence), also on a “not found”: an emptied or unmounted storage target answers exactly so. - Sample. From the items that carry content, a random sample is drawn per category: 20 mails, 20 files, 5 calendar items and 5 contacts by default. Folders, placeholders, historical file versions and parts of split messages are not drawn: they are empty or covered by their parent. The random seed is stored with the result. A requested sample size may be 1 to 200 per category; calendar and contact items get a quarter of it, rounded up. A health check reads every eligible item instead of a sample.
- Read-back through the restore path. Each item is read with the same code a restore uses: chunk index lookup, pack fetch with a fallback to a copy target, AES-256-GCM decryption. Every chunk is re-addressed twice: the id bound into its sealed header must equal the id the manifest names (checked before decrypting), and the HMAC id recomputed from the decrypted plaintext must equal it too. A chunk that decrypts cleanly but is the wrong chunk fails here.
- Whole item. The number of bytes read must equal the size in the manifest, and a fresh SHA-256 of the bytes must equal the manifest’s object hash. Items whose manifest entry carries no object hash are marked
not_recordedand rely on the chunk checks alone. - Damaged storage. Packs the latest storage check left corrupt are looked up for this snapshot through the chunk index. If the snapshot uses one, the result is red (
storage_corrupt), even when the sample did not touch that pack.
Each sampled item ends as one of:
| Item status | Meaning |
|---|---|
verified |
every chunk re-addressed, size and object hash match |
mismatch |
the bytes came back but differ from the manifest (size, object hash or a chunk that is not the one its id names) |
missing |
chunks the manifest names are not in the chunk index, or a pack that every storage target answered it does not have |
unreadable |
stored data that does not decode: a damaged pack, or a chunk that fails AES-GCM authentication or that its pack does not hold; the report names the classified cause (for example a damaged pack or a key that does not open the data) |
An item whose read failed for any other reason gets no status: it proves nothing about the backup and makes the check not completed (see Red only with evidence).
How the result is rated
Section titled “How the result is rated”The rating is the worst severity among the reasons the check found:
| Reason | Severity | When |
|---|---|---|
no_snapshot |
red | no completed backup exists |
manifest_unreadable |
red | the manifest of the latest backup came back but does not decode |
items_missing, items_unreadable, items_mismatched |
red | at least one sampled item came back missing, damaged or different (the item statuses above) |
storage_corrupt |
red | the snapshot uses a pack the storage check found damaged without an intact copy |
test_restore_failed |
red | a test restore into a target failed for a classified, lasting reason (not connected in 0.1.0, see below) |
snapshot_outdated |
red | the latest backup is 7 days (168 hours) old or older: recent data would be lost. When this is the only red reason, the bell and the alert mail say Backup too old, not Restore check failed |
snapshot_stale |
yellow | the latest backup is 48 hours old or older |
nothing_to_verify |
yellow | the snapshot holds nothing that could be read back |
test_restore_unconfirmed |
yellow | a test restore landed but the target did not confirm it (not connected in 0.1.0) |
No reason means green: every sampled item came back intact from a current backup. In the interface green is Ready, yellow Attention, red Not restorable. The report of each check shows the cause of every failed item.
Red only with evidence
Section titled “Red only with evidence”A check rates red only with evidence that the backup itself is broken, the same rule as for servers and clients:
- data that is missing: a chunk the chunk index does not know, or a pack that every storage target answered it does not have (a definite “not found”, never a timeout),
- data that does not match: a size or SHA-256 that differs, or a chunk that is not the one its id names,
- stored data that does not decode: a damaged pack, or a chunk that fails AES-GCM authentication.
Anything else proves nothing about the backup: a network error, a timeout, a 5xx answer, throttling, a reset connection, a DNS failure, and any error Restow does not recognise. Such a read makes the check not completed:
- The read-back stops at the first such read instead of waiting out the same failure for every further item.
- The check writes no report and leaves the object’s last check as it was. It raises no notification, no
verify.completedand nojob.failed. - The job is repeated with the queue’s backoff (three more attempts, from two minutes). The last attempt completes the job without a rating, and the next scheduled check tries again.
- History and the run’s page show the check as Not completed, will be retried, in a neutral tone. While attempts remain, the run’s page also explains the cause (for example the storage was unreachable or rate limited).
A copy target that answers still serves the check when the primary does not. Evidence found before the storage stopped answering still rates red, and so does a pack the storage check found damaged (storage_corrupt). A backup that is merely old is no such evidence: a check of an old backup that could not read its data is not completed too.
When it runs
Section titled “When it runs”- Weekly, by the recommended verification schedule: Sunday 03:00 in the tenant’s time zone (
Europe/Berlinunless set otherwise). - After every backup, while the tenant has an enabled verification schedule: the new snapshot is checked, unless a check of that object is already queued or running.
- On request: Check now for one object and Check all now for the tenant on the Recovery readiness page.
- Queue priority: restores before checks, checks before backups.
A rating older than 8 days counts as overdue.
The test-restore probe
Section titled “The test-restore probe”The check can, by design, also restore the verified sample into a separate test target and wait for the target to confirm each item (test_restore_failed, test_restore_unconfirmed). A failed item counts as red only for a classified, lasting reason (a permission, a missing mailbox); throttling or an unknown error makes the check not completed. In 0.1.0 nothing supplies such a target: checks perform the internal read-back only and do not restore into a Microsoft 365 mailbox, OneDrive or IMAP account.
The storage check (scrub)
Section titled “The storage check (scrub)”The storage check looks at the pack files themselves, independent of any one backup:
- Sampled, weekly (Saturday 04:00 in the recommended schedule): 5 percent of the tenant’s packs, at least 16, drawn at random, plus every pack an earlier check left corrupt or marked damaged.
- Full, monthly (the 1st, 05:00): every pack.
Each selected pack is read from every storage target that should hold it (the primary and each copy) and compared with what PostgreSQL recorded when it was written: the size, the SHA-256 of the whole file, a well-formed pack (header, index and footer, verified when the pack is opened), the tenant named in its header, and an index that still holds every chunk the chunk index places in it, at the same offset and length.
- A copy that fails while another target holds an intact copy is repaired: the intact bytes are written back under the same key and read again to confirm the SHA-256.
- A pack without an intact copy anywhere is corrupt. Every protected object whose active snapshots use it gets a red report at once (it does not wait for the next weekly check), and an alert is raised.
- A corrupt pack is marked damaged only when no target failed with an input or output error, that is, when the damage is proven and not just an unreachable target. Later backups then stop deduplicating against its chunks and write intact copies of whatever the source still holds.
Before the integrity pass, the same job brings every copy target up to the primary. Sampled runs compare copies by size, full runs by SHA-256.
Servers and clients (endpoint backup)
Section titled “Servers and clients (endpoint backup)”A machine’s backup is a restic repository on the Restow instance. restic stores its data content-addressed (SHA-256 of each blob) and encrypted (AES-256 in counter mode with Poly1305-AES authentication), see the restic design document.
What the agent records
Section titled “What the agent records”After every successful backup the agent picks sample files from the snapshot: up to 80 random regular files as candidates, none larger than 256 MiB, at most 1 GiB to hash in total. It hashes a candidate on disk only when the file is provably the one the snapshot saw: same size as in the snapshot, modification time within two seconds of the snapshot’s, and both unchanged before and after hashing. Up to 20 such files are reported with path, size and SHA-256, and stored as the backup’s samples (kept 90 days). A file edited after the backup is skipped rather than recorded with a hash the backup never had.
The server’s restore test
Section titled “The server’s restore test”The server tests every new good backup that has samples and no server-side test yet. It reads each sampled file back from the repository with restic dump, streams it through SHA-256 without writing it to disk, and compares the hash with the one the agent reported.
- The smallest file is read first and alone, the others four at a time. restic 0.19 cannot set up a new cache folder from several processes at once: one process writes the folder’s version file while another reads it still empty and gives up (“unable to open cache: readVersion”). That made the first restore test of a newly enrolled machine fail at random. Once one process has set the folder up, restic shares it safely.
- Green: every sampled file matched, and at least one file was tested.
- Red needs proof: a hash that differs, or restic ending the read with exit code 1 and, as its last line, a fatal error about the backup itself: the file is not in the snapshot, a
data,indexorsnapshotobject does not exist, an object is not found in the repository, ciphertext verification failed, or invalid data was returned. Earlier warning lines do not count, only the error restic gave up with. - Anything else is incomplete, not red: the repository busy or locked, restic stopped by a signal, crashed, unable to start, timed out, unable to reach the repository, or a wrong password or missing repository (those are the repository check’s business). The job then writes no report and fails with
RestoreTestIncompleteError; the queue retries it (3 times, from 2 minutes with backoff) and the scheduler offers it again every hour until a test completes. Until then the machine stays Not verified. A failed read never counts as a match. - The test shares the repository with browsing and downloads and never runs beside retention or the repository check (a PostgreSQL advisory lock per machine). A job that cannot get the lock within a minute is retried, without a rating.
- An administrator can start a test from the machine’s page (Run restore test).
The agent’s own restore test
Section titled “The agent’s own restore test”After a rated server test, the server hands the same files to the agent as a verify_sample task (valid seven days). The agent restores all of them with one restic restore into a temporary folder on the machine, hashes each file and deletes the copy. It sends no verdict, only what it found: per file the SHA-256 of the restored copy, missing when nothing is at that path, or an error when the machine could not check the file (not a regular file, unreadable); and, when restic restore failed, restic’s exit code, the error it gave up with and the errors it reported for single items. The server judges the report by the same rules as its own test:
- Green: the run succeeded and every file came back with its hash.
- Red needs proof: a hash that differs; a file missing from the restored copy after a
restic restorethat ended without error (the file is not in the snapshot); or restic’s own findings: it ended with exit code 1 and either the error it gave up with is one of the findings listed for the server’s test, or it gave up with “There were N errors”, the agent forwarded all N, and every item that failed has at least one such finding among its errors. “The snapshot does not exist” is no proof when the tested snapshot is no longer the machine’s newest backup: retention may have forgotten it. - Anything else is incomplete, not red: for example a full or unwritable disk on the machine, the agent stopped or restarted during the test, a busy or unreachable repository, a restic that could not start, a file the machine could not check, or findings mixed with other errors. The test then rates nothing: no report, Last restore test does not move, no alert and no
job.failed. The machine is offered the same test again after 1, 2, 4, 8, 16 and 24 hours, as long as it is not revoked, the tested backup is its newest and no test of that backup waits already; after that, the next backup brings a new test. The same applies when the monitor closes a test whose agent has not reported for six hours. The machine’s page shows such a test as Not completed, will be retried, in a neutral tone with its cause, and when the next attempt will be picked up.
A judged result becomes a second report on the same snapshot (origin agent), and a red report of either origin makes the machine red.
The repository check
Section titled “The repository check”Once a week the server runs restic check --read-data-subset=n/20 on each machine’s repository: one twentieth of the stored data, the slice advancing with the week number, so the whole repository is read back over 20 weeks (about five months) without ever reading all of it at once. restic verifies the repository structure and the hashes and authentication of the data it reads.
- A check that passes writes a green report.
- A check that could not run (the repository locked, restic interrupted or stopped by a signal, restic unable to start) is no check: nothing is reported and it is offered again. Six locked attempts in a row over at least twelve hours raise the alert
endpoint.repository_locked. - Any other restic error writes a red report with restic’s message.
How a machine is rated
Section titled “How a machine is rated”The newest restore test of each origin (server, agent) for the machine’s newest backup counts:
| State | When |
|---|---|
| No backup yet | no good backup exists |
| Not verified | a backup exists, but no completed restore test of exactly that backup |
| Ready | the restore tests of the newest backup are green |
| Attention | green, but the backup itself was partial (some files could not be read on the machine) |
| Not restorable | a restore test of the newest backup is red, or a repository check newer than the last test found damage |
A rating older than 8 days counts as overdue.
Evidence chains
Section titled “Evidence chains”Two hash chains make later changes to records visible. They are not restore checks, but they are part of what a report can prove:
- Audit log. Each entry’s hash is SHA-256 over the predecessor’s hash, the entry’s fields as canonical JSON and its timestamp, one chain per tenant and one for the installation. The database roles of the application cannot change or delete entries, and a daily seal of each chain’s last entry makes a cut-off or recomputed chain detectable. Verifying the chain in the interface is Business and Service Provider; recording happens in every edition.
- Archive. Each archived item’s chain hash is SHA-256 over the predecessor’s chain hash, the SHA-256 of the archived original and the time of receipt. The archive page verifies the chain.
Where results are stored and shown
Section titled “Where results are stored and shown”| What | Stored in | Shown |
|---|---|---|
| Restore check of a mailbox, OneDrive or IMAP account | verify_reports (rating, snapshot id, reasons, counts, the sampled items, the seed, durations); the job record |
Recovery readiness (table, report per object), Overview (tabs Status and Statistics), History |
| Storage check | the job record; a red verify_reports row per affected object |
Recovery readiness, Repositories page, alerts |
| Restore test and repository check of a machine | endpoint_reports (kind, origin, snapshot id, rating, summary with the mismatched files) |
Recovery readiness, the machine’s page |
| Rating changes | notifications verify.red (worded Backup too old when the only red reason is an old backup), verify.yellow, verify.recovered; for damaged storage scrub.corrupt |
the bell, alert rules (mail, webhook), the delivery log |
| Every rated mailbox check (not one that could not complete) | webhook event verify.completed |
your RMM or PSA |
The integration API reports the same states for automation.
Verifying a backup yourself
Section titled “Verifying a backup yourself”The checks above run inside Restow. To check independently of a running Restow server:
- Without the server or its database:
restow-restore verifyreads a snapshot from the storage with the master key (or an exported keyring) and runs the same reassembly a restore does, writing nothing: every chunk must decrypt and carry the id the manifest names, and every item must have the size the manifest records.restow-restore restorewrites the files out. Both exit non-zero if any object fails. See Backing up Restow itself and Verifying a release. - A machine’s repository: open its repository password with
restow-restore endpoint-password(master key and storage are enough), then use plain restic, for examplerestic check --read-dataor a restore into a test folder. See Endpoint backup: restore. - End to end: restore a test mailbox, a OneDrive folder or a machine’s folder from time to time and look at the result. The automated checks do not write into a live mailbox or onto a machine’s existing files.
Limits
Section titled “Limits”- A sample is a sample. 20 mails of a large mailbox prove that the restore path and those mails work, not that every other mail does. The storage check complements it by hashing every pack each month; a health check reads every item when you ask for it.
- The checks prove that the backup on your storage can be read back intact. They do not prove that writing it into a live Microsoft 365 mailbox, OneDrive or IMAP account succeeds today; the test-restore probe is not connected in 0.1.0.
- Microsoft 365 backup and restore have not yet run against a real Microsoft 365 tenant in the release checks; they are covered by tests against a simulated Graph API.
- A machine’s sample covers files that were unchanged during the backup. Whether a database dump or an open file is consistent is up to the hooks that prepare it.
- A storage target that loses its mount in the middle of a check, after the manifest was read, answers “not found” for the data files, and that counts as missing data: the object is rated red. A target that is unmounted before the check makes it not completed.
- The checks detect damage, they do not prevent deletion: protect the storage itself (see the operator notice).
Implementation
Section titled “Implementation”The behaviour above, in the product’s source (repository restow-backup/restow):
- Mailbox restore check:
packages/core/src/verify/engine.ts,check.ts,evidence.ts(what counts as evidence),errors.ts(VerifyIncompleteError),sampling.ts,readiness.ts,report.ts; the worker jobapps/worker/src/handlers/verify.ts; the rule that a rating belongs to its snapshotapps/api/src/features/verify/verification-state.ts. - Chunks, packs and manifests:
packages/core/src/crypto.ts,chunkId.ts,pack.ts,manifest.ts,engine/chunkstore.ts(openChunkAs,RestoreIntegrityError). - Storage check:
packages/core/src/verify/scrub.ts,integrity.ts; the worker jobapps/worker/src/handlers/scrub.ts. - Schedules:
packages/core/src/schedule/defaults.ts; endpoint jobsapps/scheduler/src/endpoints.ts, queue settingspackages/core/src/endpoints/queues.ts. - Endpoint restore test:
packages/core/src/endpoints/restore-test.ts(restoreTestSamples,isBackupFinding,isRestoreFinding,judgeAgentRestoreTest), the waits of a repeated testpackages/core/src/endpoints/restore-test-retry.ts, the jobapps/worker/src/endpoints/verify.ts(RestoreTestIncompleteError), the monitorapps/worker/src/endpoints/monitor.ts, the agent’s samplesagent/internal/core/sample.go, the agent’s testagent/internal/core/verify.go, its reportapps/api/src/features/endpoints/agent-service.ts. - Repository check:
apps/worker/src/endpoints/check.ts; ratingpackages/core/src/endpoints/readiness.ts. - Evidence chains:
packages/core/src/audit-chain.ts,packages/core/src/archive/chain.ts. - Standalone tool:
packages/cli.