| 1 | # Testing — Why the code can be trusted |
| 2 | |
| 3 | Snowbound writes into other people's notebooks, often while their copy of |
| 4 | OneNote is writing to the same file. A bug doesn't just crash an app. It can |
| 5 | corrupt a shared notebook, or silently drop a paragraph someone else typed. |
| 6 | The testing strategy follows from that. **Never grade the code with itself.** |
| 7 | Every important claim is checked against something the code under test didn't |
| 8 | produce: a real OneNote, a separate model, a second path through the system, |
| 9 | or a disk that loses bytes on purpose. |
| 10 | |
| 11 | This essay explains where the confidence comes from. The how-to for running |
| 12 | each lane is in [tools/TESTING.md](../tools/TESTING.md) and the crate READMEs. |
| 13 | |
| 14 | ## The reference is a real OneNote |
| 15 | |
| 16 | The only authority on "OneNote accepts this" is OneNote. The lab (`tools/w7`) |
| 17 | runs OneNote 2010 in disposable QEMU clones of a sealed Windows 7 image. An |
| 18 | agent inside each clone runs AutoHotkey and PowerShell (OneNote's COM API), |
| 19 | returns screenshots and files, and the clone is discarded afterwards. |
| 20 | |
| 21 | Every storage feature goes through the same loop: |
| 22 | |
| 23 | ```text |
| 24 | observe author the edit in OneNote (COM, or driving its UI) and dump what it stored |
| 25 | write make Snowbound store the same thing through ops; export a candidate notebook |
| 26 | cold-open a fresh clone with a fresh OneNote cache opens the candidate |
| 27 | ─► integrity check passes, XML export, screenshots, PDF where layout matters |
| 28 | gate a VM-free test compares the capture with what the candidate meant to say |
| 29 | install candidate + capture become a corpus row (corpus/<feature>/…) |
| 30 | ``` |
| 31 | |
| 32 | *Cold* matters. A warm OneNote has its own cached copy and can hide a file it |
| 33 | would reject. Some failures appear only at scale. OneNote once silently dropped |
| 34 | elements in about one build in five because of an identity choice its integrity |
| 35 | check accepted, and it refuses revision chains past a certain depth. Gates |
| 36 | therefore include large mixed candidates, not only one small row per feature. |
| 37 | |
| 38 | The corpus is what makes this sustainable. Each row keeps the native capture |
| 39 | next to the candidate, and its `tools/test_*.py` gate checks it without a VM, |
| 40 | on every run, forever. Re-running the VM step is needed only when the bytes |
| 41 | Snowbound writes change. |
| 42 | |
| 43 | ## Oracles inside the build |
| 44 | |
| 45 | Most checking needs no VM. It works by making independent views of one edit |
| 46 | agree: |
| 47 | |
| 48 | - **The model oracle.** `onestore::op::model` interprets ops on the page model |
| 49 | without going near bytes. After a random edit is applied to a `Section`, |
| 50 | sealed and read back, the page must equal what the model predicted. |
| 51 | Refused edits must leave the page unchanged. |
| 52 | - **Differential editing.** For every editor action, alone and in random |
| 53 | sequences with undo and redo, on every page of a set of corpus sections, |
| 54 | four views must agree: the page stored through the ops, the model's |
| 55 | prediction from those ops, the sealed image read back, and the editor's own |
| 56 | page (`crates/canvas/tests/ops_differential.rs`). |
| 57 | - **Seals check themselves.** Each seal validates exactly what it appended. The |
| 58 | full-file validator runs on open and over the tests' images. |
| 59 | - **Layout agrees with itself and with OneNote.** Incremental layout must |
| 60 | equal a fresh layout after long structural histories. Undo and redo must |
| 61 | restore geometry exactly. Unicode range replacement is compared with plain |
| 62 | string replacement across hundreds of combinations. Wraps and heights are |
| 63 | compared with OneNote's own XML and PDF exports. |
| 64 | - **Conversions verify their output.** A cache schema migration must reproduce |
| 65 | every page the old cache held, or it rolls back. |
| 66 | |
| 67 | ## Failure, on purpose |
| 68 | |
| 69 | Durability claims are tested by making things fail at every point that |
| 70 | matters. |
| 71 | |
| 72 | - **Torn writes.** The storage crash model is a disk that, at an injected |
| 73 | failure, persists a random subset of the bytes written since the last flush. |
| 74 | Every I/O point of ordinary commits and counter-rollover commits is cut. |
| 75 | Afterwards the file must read as the old image or the new one, never a |
| 76 | third thing. The same model drives the commit fuzzer. |
| 77 | - **Crashes in the replica.** SQLite transactions are cut while their frames |
| 78 | are in the log but the commit frame isn't, at every step of edit, seal, |
| 79 | publish and acknowledge. Reopening must recover the queue as it was. |
| 80 | - **Lost replies.** `tools/smb-proxy.py` sits between the client and a lab |
| 81 | Samba server and withholds a chosen request or response. The matrix cuts |
| 82 | every message of a commit. Each outcome must keep its promise: |
| 83 | `NotCommitted` really didn't publish, `Committed` really did, and `Unknown` |
| 84 | cases really went either way, with the confirmation that follows finding |
| 85 | out which. |
| 86 | - **Power cuts.** The lab can kill a VM outright mid-workload, and the |
| 87 | surviving files are then cold-opened by a real OneNote. |
| 88 | - **Schedules.** Deterministic multi-actor tests run several replicas editing |
| 89 | offline and publishing through a fault-injecting in-memory remote, then |
| 90 | check convergence and conflict pages. |
| 91 | |
| 92 | ## Real clients on a real share |
| 93 | |
| 94 | The collaboration lab puts several OneNote 2010 clients and several Snowbound |
| 95 | writers and readers on one section in a Samba share. They make random edits |
| 96 | through outages, reconnects, offline stretches and OneNote's own maintenance |
| 97 | (Optimize rewrites the file underneath everyone). The pass criteria are strict: |
| 98 | |
| 99 | - every intent from every client is in the final file exactly once, or is |
| 100 | accounted for as a conflict page; |
| 101 | - a wire trace shows OneNote was never refused a read; the only refusals it |
| 102 | met are the ones its own protocol makes between two writers; |
| 103 | - a fresh OneNote cold-opens the result and agrees. |
| 104 | |
| 105 | ## Fuzzing |
| 106 | |
| 107 | `fuzz/` is its own cargo-fuzz workspace. It covers parsing (storage, property |
| 108 | streams, revisions, documents), the commit protocol with interruptions, edits |
| 109 | of many kinds, the page model, protected sections, the offline queue, and the |
| 110 | canvas editor's state machine. The parsers see arbitrary bytes, and the |
| 111 | writers see arbitrary sequences of edits whose results must still validate. |
| 112 | Clipboard text and clips, Live Share messages and SMB directory records have |
| 113 | targets too. The SMB target compiles the directory decoder's source directly, |
| 114 | without a connection or a separate parser. |
| 115 | |
| 116 | ## Two tiers of tests |
| 117 | |
| 118 | Tests fall into two kinds, and the suite is being split to match: |
| 119 | |
| 120 | - **Correctness tests** are deterministic and run on every change: `cargo test` |
| 121 | over the workspace, clippy, and the Python gates over retained captures. |
| 122 | `tools/check_public.py` runs all of it from a clean checkout, with no private |
| 123 | notebooks, VMs or credentials. |
| 124 | - **Sweeps and fuzzing** are random sequences over large corpora and |
| 125 | coverage-guided fuzzers. They are opt-in (today, sweeps widen through |
| 126 | environment variables such as `CANVAS_SWEEP_*` and `OPS_SWEEP_SECTIONS`) and |
| 127 | are meant to run for days or weeks on dedicated machines. What they find gets |
| 128 | reduced to a small deterministic test in the first tier. |
| 129 | |
| 130 | Lab lanes (real OneNote, Samba VMs, the proxy) sit beside both tiers. In Rust |
| 131 | they appear as ignored tests that name the environment they need, and the rest |
| 132 | live in Python harnesses under `tools/`. |
| 133 | |
| 134 | ## What counts as evidence |
| 135 | |
| 136 | The project is strict about what a result establishes, and that strictness is |
| 137 | itself a source of confidence: |
| 138 | |
| 139 | - An ignored test, a missing capture, an unreachable lab, a mocked harness or |
| 140 | an iOS build that only compiled establishes nothing about native |
| 141 | compatibility. |
| 142 | - A process exit, a VM interruption and a physical power loss are different |
| 143 | fault models, and one never stands in for another. The evidence covers |
| 144 | transport failures and VM cuts, not every server's lock implementation or |
| 145 | real hardware losing power. |
| 146 | - Notebooks under test are always disposable copies with recorded hashes. |
| 147 | Nothing points an editing harness, a replay or the app at an original. |
| 148 | - Accessibility is tested by asserting on the AccessKit tree the app builds |
| 149 | and on the running app's accessibility hierarchy. Nobody turns a screen |
| 150 | reader on to listen for it. |