1# Testing — Why the code can be trusted
2
3Snowbound writes into other people's notebooks, often while their copy of
4OneNote is writing to the same file. A bug doesn't just crash an app. It can
5corrupt a shared notebook, or silently drop a paragraph someone else typed.
6The testing strategy follows from that. **Never grade the code with itself.**
7Every important claim is checked against something the code under test didn't
8produce: a real OneNote, a separate model, a second path through the system,
9or a disk that loses bytes on purpose.
10
11This essay explains where the confidence comes from. The how-to for running
12each lane is in [tools/TESTING.md](../tools/TESTING.md) and the crate READMEs.
13
14## The reference is a real OneNote
15
16The only authority on "OneNote accepts this" is OneNote. The lab (`tools/w7`)
17runs OneNote 2010 in disposable QEMU clones of a sealed Windows 7 image. An
18agent inside each clone runs AutoHotkey and PowerShell (OneNote's COM API),
19returns screenshots and files, and the clone is discarded afterwards.
20
21Every storage feature goes through the same loop:
22
23```text
24observe author the edit in OneNote (COM, or driving its UI) and dump what it stored
25write make Snowbound store the same thing through ops; export a candidate notebook
26cold-open a fresh clone with a fresh OneNote cache opens the candidate
27 ─► integrity check passes, XML export, screenshots, PDF where layout matters
28gate a VM-free test compares the capture with what the candidate meant to say
29install candidate + capture become a corpus row (corpus/<feature>/…)
30```
31
32*Cold* matters. A warm OneNote has its own cached copy and can hide a file it
33would reject. Some failures appear only at scale. OneNote once silently dropped
34elements in about one build in five because of an identity choice its integrity
35check accepted, and it refuses revision chains past a certain depth. Gates
36therefore include large mixed candidates, not only one small row per feature.
37
38The corpus is what makes this sustainable. Each row keeps the native capture
39next to the candidate, and its `tools/test_*.py` gate checks it without a VM,
40on every run, forever. Re-running the VM step is needed only when the bytes
41Snowbound writes change.
42
43## Oracles inside the build
44
45Most checking needs no VM. It works by making independent views of one edit
46agree:
47
48- **The model oracle.** `onestore::op::model` interprets ops on the page model
49 without going near bytes. After a random edit is applied to a `Section`,
50 sealed and read back, the page must equal what the model predicted.
51 Refused edits must leave the page unchanged.
52- **Differential editing.** For every editor action, alone and in random
53 sequences with undo and redo, on every page of a set of corpus sections,
54 four views must agree: the page stored through the ops, the model's
55 prediction from those ops, the sealed image read back, and the editor's own
56 page (`crates/canvas/tests/ops_differential.rs`).
57- **Seals check themselves.** Each seal validates exactly what it appended. The
58 full-file validator runs on open and over the tests' images.
59- **Layout agrees with itself and with OneNote.** Incremental layout must
60 equal a fresh layout after long structural histories. Undo and redo must
61 restore geometry exactly. Unicode range replacement is compared with plain
62 string replacement across hundreds of combinations. Wraps and heights are
63 compared with OneNote's own XML and PDF exports.
64- **Conversions verify their output.** A cache schema migration must reproduce
65 every page the old cache held, or it rolls back.
66
67## Failure, on purpose
68
69Durability claims are tested by making things fail at every point that
70matters.
71
72- **Torn writes.** The storage crash model is a disk that, at an injected
73 failure, persists a random subset of the bytes written since the last flush.
74 Every I/O point of ordinary commits and counter-rollover commits is cut.
75 Afterwards the file must read as the old image or the new one, never a
76 third thing. The same model drives the commit fuzzer.
77- **Crashes in the replica.** SQLite transactions are cut while their frames
78 are in the log but the commit frame isn't, at every step of edit, seal,
79 publish and acknowledge. Reopening must recover the queue as it was.
80- **Lost replies.** `tools/smb-proxy.py` sits between the client and a lab
81 Samba server and withholds a chosen request or response. The matrix cuts
82 every message of a commit. Each outcome must keep its promise:
83 `NotCommitted` really didn't publish, `Committed` really did, and `Unknown`
84 cases really went either way, with the confirmation that follows finding
85 out which.
86- **Power cuts.** The lab can kill a VM outright mid-workload, and the
87 surviving files are then cold-opened by a real OneNote.
88- **Schedules.** Deterministic multi-actor tests run several replicas editing
89 offline and publishing through a fault-injecting in-memory remote, then
90 check convergence and conflict pages.
91
92## Real clients on a real share
93
94The collaboration lab puts several OneNote 2010 clients and several Snowbound
95writers and readers on one section in a Samba share. They make random edits
96through outages, reconnects, offline stretches and OneNote's own maintenance
97(Optimize rewrites the file underneath everyone). The pass criteria are strict:
98
99- every intent from every client is in the final file exactly once, or is
100 accounted for as a conflict page;
101- a wire trace shows OneNote was never refused a read; the only refusals it
102 met are the ones its own protocol makes between two writers;
103- a fresh OneNote cold-opens the result and agrees.
104
105## Fuzzing
106
107`fuzz/` is its own cargo-fuzz workspace. It covers parsing (storage, property
108streams, revisions, documents), the commit protocol with interruptions, edits
109of many kinds, the page model, protected sections, the offline queue, and the
110canvas editor's state machine. The parsers see arbitrary bytes, and the
111writers see arbitrary sequences of edits whose results must still validate.
112Clipboard text and clips, Live Share messages and SMB directory records have
113targets too. The SMB target compiles the directory decoder's source directly,
114without a connection or a separate parser.
115
116## Two tiers of tests
117
118Tests fall into two kinds, and the suite is being split to match:
119
120- **Correctness tests** are deterministic and run on every change: `cargo test`
121 over the workspace, clippy, and the Python gates over retained captures.
122 `tools/check_public.py` runs all of it from a clean checkout, with no private
123 notebooks, VMs or credentials.
124- **Sweeps and fuzzing** are random sequences over large corpora and
125 coverage-guided fuzzers. They are opt-in (today, sweeps widen through
126 environment variables such as `CANVAS_SWEEP_*` and `OPS_SWEEP_SECTIONS`) and
127 are meant to run for days or weeks on dedicated machines. What they find gets
128 reduced to a small deterministic test in the first tier.
129
130Lab lanes (real OneNote, Samba VMs, the proxy) sit beside both tiers. In Rust
131they appear as ignored tests that name the environment they need, and the rest
132live in Python harnesses under `tools/`.
133
134## What counts as evidence
135
136The project is strict about what a result establishes, and that strictness is
137itself a source of confidence:
138
139- An ignored test, a missing capture, an unreachable lab, a mocked harness or
140 an iOS build that only compiled establishes nothing about native
141 compatibility.
142- A process exit, a VM interruption and a physical power loss are different
143 fault models, and one never stands in for another. The evidence covers
144 transport failures and VM cuts, not every server's lock implementation or
145 real hardware losing power.
146- Notebooks under test are always disposable copies with recorded hashes.
147 Nothing points an editing harness, a replay or the app at an original.
148- Accessibility is tested by asserting on the AccessKit tree the app builds
149 and on the running app's accessibility hierarchy. Nobody turns a screen
150 reader on to listen for it.