Every memory product in this market publishes a number. Almost none publishes a way of arriving at one. This page is the way: what was sealed and when, what the harness does, what it has measured so far, and what it has not. The quiz stays sealed, and the last section says why that is the point rather than the excuse.
Vendors in this field quote accuracy in the nineties on the public memory benchmarks. Every independent rerun this year lands between the low fifties and the mid seventies. Nobody, vendor or academic, publishes what a query cost in tokens or in money, and one paper put a rival's write path at fifty times the cost of plain retrieval. When the gap between a claim and a rerun is thirty points, the claim carries no information, and a reader is right to ignore it.
On a sealed, pre-registered quiz of 120 questions over 1,246 cards written to be messy, Kit retrieval answered 60 under the custodian's rulings against 56 for plain retrieval over the same cards and 43 for the named graph rival, and found the right note 99 times to the rival's 48.
round 4 · ruled claims · 8 September 2026It is one operator's corpus, written by that operator's own agents to a design the builder wrote, run by the system under test. Kit's lead over plain retrieval on the ruled column is not settled at this sample; its lead over the rival is. Any of those is enough to withhold belief.
the cold review, 2 September 2026, still standingThe checksums, method and summary results are displayed here. A checksum alone does not give a visitor the underlying corpus or harness. The interval, so the sample is visible. The cost, so the bill is not hidden. Then the failures beside the successes.
eval/bakeoff, in this repositoryTwo files carry the seal. One freezes the corpus every contestant answered over. The other freezes the study's own rules: the hypothesis, the pass bars, the disqualifiers and the pinned settings, all written down before a single answer was graded. Both are plain SHA-256 files, so checking them takes one command each and no trust in anybody.
Four story-card files, 286 cards, sealed 27 July 2026 at 08:34:31 UTC. Every contestant is given these bytes and no others. Re-exporting or editing a card breaks the seal, which is the whole reason the seal is a file rather than a promise.
Registered 27 July 2026 at 08:04:46 UTC, half an hour before the corpora were sealed, and amended once at 08:35:00 UTC to record the seal itself. It names the hypothesis, the five contestants, the six pass bars, the seven disqualifiers, and the pins the run answers at.
Eight story-card files, 1246 cards, sealed 7 September 2026 at 15:05:41 UTC: the four July files, untouched, and four new volumes written to be messy the way a real store is. Every contestant is given these bytes and no others.
Registered 7 September 2026 at 15:21:54 UTC, after the corpora were sealed and before the quiz was written. It names the hypothesis, the five contestants, the seven question classes, the three key policies the quiz was written to, the pins the run answers at, and the hash of the sealed quiz and key.
# from the repository root shasum -a 256 -c eval/bakeoff/CORPORA_SEAL.sha256 shasum -a 256 -c eval/bakeoff/ROUND4_SEAL.sha256 # the registration names its file by basename, # so it is checked from the folder it sits in cd eval/bakeoff && shasum -a 256 -c WEEK0_REGISTRATION.sha256 # round 4's registration names its file from the repository root cd .. && shasum -a 256 -c eval/bakeoff/ROUND4_REGISTRATION.sha256
On Linux the command is sha256sum -c and the arguments are the same. Fourteen lines come back, each ending OK. One line that does not is a broken seal, and a broken seal invalidates every number below it.
Five readers answer the same questions over the same sealed cards under the same model, the same budget and the same instructions. Only the memory system behind the reader changes. Everything the run computes is derived from the rows it recorded, and a second program re-derives it from the file afterwards, so a figure in a published artifact is checkable by whoever holds the artifact.
A no memory. B plain hybrid retrieval over the same cards, which the registration calls the honest baseline. C Kit retrieval only. D full Kit, with supersession and cite-or-abstain. E Graphiti, the named graph rival.
readers/contestants.pyPrimary is claims: required claims present, forbidden claims absent, superseded guidance not repeated, and an abstention where the corpus genuinely has no answer. Secondary is cites, which also requires a resolvable source id. Both are always reported.
honesty.py · seven labelsA system that says it does not know scores on that question. A system that invents an answer loses on it. Seven labels separate a correct abstention from a missed one and an invention from a stale answer, so honesty is graded rather than assumed.
supported_correct, invent, stale, scope_leak, abstain_correct, abstain_miss, unsupportedEvery answer carries its model, its tokens and its price, from one rate table and one rule. The reported figure for the Kit contestants is what the ledger charged, read back, not a second estimate of it.
readers/cost.pyDollars per card on the way in, measured on both sides or on neither. One side metered is worse than none, because the asymmetry is the finding.
readers/write_cost.pyAn optional second pass where a different frontier model sees the question, the rubric and one answer, blind to which contestant produced it, and replies pass or fail. It corroborates; it never replaces. Disagreements are listed, not resolved.
judge.pyThe answering instructions were a hidden variable: three copies in the tree, and the copy the Kit readers answered under carried clauses the baseline never saw. There is now one copy and a named profile, and a neutral rerun writes its own artifact.
readers/answer_prompt.pyEvery rate carries a 95% Wilson interval, and every gap against the baseline is tested paired on the questions both answered, exactly rather than through an approximation that wants far more disagreements than this quiz produces.
margin.pyThe whole quiz can be asked N times over, reporting the range beside the mean and naming the questions whose verdict moved. The temperature comes out of the seal rather than out of the code, and every row records the one it answered at.
repeats.py · readers/temperature.pyRows are the evidence and every other figure is derived, so the derived figures are recomputed from the rows and compared. A run audits the file it has just written and exits non-zero when the two disagree.
report.pyNo database, no container, no answer key, no network, no API key. It reads one file in this repository and rebuilds every figure in it from the rows underneath. The artifact below is the apparatus run, which is Kit-authored and deliberately not the sealed study, and it can be checked by someone with access to those private repository files. It is not downloadable from this page.
python3 eval/bakeoff/report.py eval/bakeoff/mechanism/last_differentiated_run.json B Plain RAG: claims 15/20 | 95% CI 53.1% to 88.8% D Full Kit: claims 18/20 | 95% CI 69.9% to 97.2% D Full Kit: +15.0 points over 20 questions | won 3, lost 0, tied 17 | exact p 0.250 clears the registered 10 point bar, and the paired test cannot tell this gap from chance about 37 questions at this rate would settle it (the study registered 120)
That last line is the harness reporting against its own product, unprompted, and it is the reason this page exists. --check prints the verdict alone and sets the exit code, so the check is a command rather than a reading exercise.
Round 4 of the custodian's quiz: 120 questions in seven classes over the 1,246 sealed cards, answered under the tuned profile at the registered temperature, eight notes for every contestant, by Kit as it ships on 8 September (main at 000fafa5). The column to quote is the ruled one: the strict key's verdict, corrected only on the rows where the key and the blind judge disagreed and only by the five written policies, with 94 rows ruled and 1 left open. The key and the judge stand beside it, and the interval sits beside the rate rather than underneath it.
| contestant | ruled | rate | 95% interval | key | judge | found the note | read cost | write cost |
|---|---|---|---|---|---|---|---|---|
| C · Kit retrieval | 60 / 120 | 50.0% | 41.2% to 58.8% | 36 | 61 | 99 | $1.11 per run | $0.0018 per card |
| D · full Kit | 61 / 120 | 50.8% | 42.0% to 59.6% | 40 | 61 | 99 | $1.12 per run | $0.0018 per card |
| B · plain hybrid retrieval | 56 / 120 | 46.7% | 38.0% to 55.6% | 38 | 54 | 95 | $1.05 per run | a local embedding, no provider call |
| E · Graphiti | 43 / 120 | 35.8% | 27.8% to 44.7% | 37 | 42 | 48 | $1.44 per run | $0.051 per card at list |
| A · no memory | 30 / 120 | 25.0% | 18.1% to 33.4% | 30 | 32 | 30 | nothing, by construction | nothing, by construction |
Found the note: how often the right card was among the eight shown, out of 120. Read cost is what the ledger charged for one pass over the quiz, answering and marking excluded. Write cost is what putting the corpus in cost per card, measured on both sides: Kit $2.19 for 1,246 cards, all of it the conflict judge; Graphiti $58.11 at list price for the 960 cards its graphs lacked, over 3 attempts under a $100 cap. Per card the rival's write path is about 29 times Kit's at list price, 34 times by the rate table.
The same quiz under the neutral profile, a plain answering prompt no contestant was coached under. The honest headline is this table beside the one above, not whichever flatters. 131 rows ruled, 1 left open.
| contestant | ruled | rate | 95% interval | key | judge | found the note |
|---|---|---|---|---|---|---|
| C · Kit retrieval | 67 / 120 | 55.8% | 46.9% to 64.4% | 38 | 64 | 99 |
| D · full Kit | 65 / 120 | 54.2% | 45.3% to 62.8% | 38 | 60 | 99 |
| B · plain hybrid retrieval | 59 / 120 | 49.2% | 40.4% to 58.0% | 21 | 57 | 95 |
| E · Graphiti | 40 / 120 | 33.3% | 25.5% to 42.2% | 36 | 42 | 48 |
| A · no memory | 30 / 120 | 25.0% | 18.1% to 33.4% | 30 | 30 | 30 |
Round 4 was built because rounds 1 to 3 could not separate Kit from a search box: 286 tidy cards, and nearly every question answerable from one card that shared its words. The new exam is messy on purpose, and Kit as it was found fewer of its own notes than plain retrieval did. Each fix below is one labelled run against the same sealed quiz, so a reader can see what each one did on its own. Key, judge and found, out of 120.
The lesson of 4e is the one worth keeping: a wider conflict net made Kit worse, because the resolver that merges two disagreeing notes was dropping their labels, so a merged note could not be found under its case file and the notes folded into it were gone. 4f fixed that and merged freely; 4g, the code that ships, merges only where it is sure, holds the rest open as uncertainty, forgets nothing and pages nobody, and gives back a few answers on the questions that need two months' notes read together. The key moves less than the judge across the series because it grades literal phrases and a merged note restates its sources; that is recorded as the key's problem, not the answer's.
The custodian's questions and answer keys are not published, and there is no plan to publish them. That decision was taken deliberately, with the cost understood, and stating the cost is part of taking it honestly.