It takes one person's own exports — a glucose sensor, a sleep ring, a chat log — answers the question they actually asked, and then spends most of its effort trying to prove that answer wrong before showing it to them.
Running the tests is the easy half. Refusing to believe them is the product.
The question is written down, with its outcome and its exposure, before the data is opened. That is what separates one confirmatory test from the hundreds of exploratory ones a dataset will happily offer you.
Empty fields in live templates, dead checkboxes, coverage gaps — and informative missingness, because people log least on their worst days, and every complete-case test quietly inherits that tilt.
Surviving a significance threshold earns a finding nothing but a turn on the stand. Seven passes then try to kill it: drop a month, reverse the arrow, adjust for the obvious third cause, ask whether the exposure is just what a bad day looks like.
HOLDS, FRAGILE, CONFOUNDED, ARTIFACT or NULL — assigned by rule rather than by taste, and reported the same way whether or not the answer is the interesting one.
One person, one question, every number recomputed from the raw export each time the page is built — so the prose cannot drift from the result it describes.
If you track something about yourself and have a real question about it, say what you track and what you want to know. A person will reply.