Kit's recall claim, and how to check it

Every memory product in this market publishes a number. Almost none publishes a way of arriving at one. This page is the way: what was sealed and when, what the harness does, what it has measured so far, and what it has not. The quiz stays sealed, and the last section says why that is the point rather than the excuse.

drawn from the harness itself · eval/bakeoff · WEEK0_REGISTRATION.json · CORPORA_SEAL.json · METHODOLOGY.txt
checked · a command in this page proves it stands · measured once, and the work behind it is unfinished not measured · instrumented, no number yet
1

Why the number is the least interesting part

Vendors in this field quote accuracy in the nineties on the public memory benchmarks. Every independent rerun this year lands between the low fifties and the mid seventies. Nobody, vendor or academic, publishes what a query cost in tokens or in money, and one paper put a rival's write path at fifty times the cost of plain retrieval. When the gap between a claim and a rerun is thirty points, the claim carries no information, and a reader is right to ignore it.

What is claimed

On a sealed, pre-registered quiz of 120 questions over 1,246 cards written to be messy, Kit retrieval answered 60 under the custodian's rulings against 56 for plain retrieval over the same cards and 43 for the named graph rival, and found the right note 99 times to the rival's 48.

round 4 · ruled claims · 8 September 2026

What is wrong with it

It is one operator's corpus, written by that operator's own agents to a design the builder wrote, run by the system under test. Kit's lead over plain retrieval on the ruled column is not settled at this sample; its lead over the rival is. Any of those is enough to withhold belief.

the cold review, 2 September 2026, still standing

What is published instead

The checksums, method and summary results are displayed here. A checksum alone does not give a visitor the underlying corpus or harness. The interval, so the sample is visible. The cost, so the bill is not hidden. Then the failures beside the successes.

eval/bakeoff, in this repository
The standard this page holds itself to: the pre-registration lists "Hiding failures, costs, or abstentions" as a disqualifier, and it was written before any of these numbers existed. Where something is unmeasured, this page says unmeasured rather than leaving the column blank, because a blank column reads as a zero and a zero is a measurement.
Public access correction, 9 September 2026: the study described here is in Kit's private repository. The commands below require its files. The questions remain sealed, and this page alone cannot reproduce the headline. The public 17-question exam is a separate, smaller test. Read the evaluation guide before treating either as evidence for your workflow.
2

The seal · what was frozen, and what a checksum proves

Two files carry the seal. One freezes the corpus every contestant answered over. The other freezes the study's own rules: the hypothesis, the pass bars, the disqualifiers and the pinned settings, all written down before a single answer was graded. Both are plain SHA-256 files, so checking them takes one command each and no trust in anybody.

The corpora · eval/bakeoff/CORPORA_SEAL.sha256

Four story-card files, 286 cards, sealed 27 July 2026 at 08:34:31 UTC. Every contestant is given these bytes and no others. Re-exporting or editing a card breaks the seal, which is the whole reason the seal is a file rather than a promise.

83122089b51614edee54ef34ed8ec6fbf647cc35be119ba4c4719395987d24d5
eval/bakeoff/papers/story_cards.json · CF1 Mayfur Society Papers · 115 cards
6be64f4c5fe7fefaf1cf9460fa58811ba0fe64c6cffa5dfca64f558cd7f0c394
eval/bakeoff/mews/story_cards.json · CF2 The Mews · 85 cards
2cbe2060befb27faf154b7efae72be5af2a6dd54920bc861feda5c837bbc54fa
eval/bakeoff/ballast/story_cards.json · CF3 Ballast case file · 62 cards
1b81d416c1f9b3a78eae23c1164b9a5b46f397a9604802f4f1486f36ab3ae165
eval/bakeoff/copperline/story_cards.json · CF4 Copperline Transit · 24 cards

The pre-registration · eval/bakeoff/WEEK0_REGISTRATION.sha256

Registered 27 July 2026 at 08:04:46 UTC, half an hour before the corpora were sealed, and amended once at 08:35:00 UTC to record the seal itself. It names the hypothesis, the five contestants, the six pass bars, the seven disqualifiers, and the pins the run answers at.

e6472c4ceead3007b949a13b2ec878c2b695f206eea2f28f6889234a9af13349
WEEK0_REGISTRATION.json · the study's rules, written before the answers

The round 4 corpora · eval/bakeoff/ROUND4_SEAL.sha256

Eight story-card files, 1246 cards, sealed 7 September 2026 at 15:05:41 UTC: the four July files, untouched, and four new volumes written to be messy the way a real store is. Every contestant is given these bytes and no others.

83122089b51614edee54ef34ed8ec6fbf647cc35be119ba4c4719395987d24d5
eval/bakeoff/papers/story_cards.json · CF1 Mayfur Society Papers · 115 cards
cc2b344e373277c8b0f1743dca2785f73d46e3a826a501d9a906b3cd746275e2
eval/bakeoff/papers-ii/story_cards.json · CF1 Mayfur Society Papers, volume two · 240 cards
6be64f4c5fe7fefaf1cf9460fa58811ba0fe64c6cffa5dfca64f558cd7f0c394
eval/bakeoff/mews/story_cards.json · CF2 The Mews · 85 cards
4c3d549e34a2db251bda9ba907140047613c190dd4b825e04f0d7da58b81a5a0
eval/bakeoff/mews-ii/story_cards.json · CF2 The Mews, volume two · 200 cards
2cbe2060befb27faf154b7efae72be5af2a6dd54920bc861feda5c837bbc54fa
eval/bakeoff/ballast/story_cards.json · CF3 Ballast case file · 62 cards
1b81d416c1f9b3a78eae23c1164b9a5b46f397a9604802f4f1486f36ab3ae165
eval/bakeoff/copperline/story_cards.json · CF4 Copperline Transit · 24 cards
5b38ddd738577bee2c5a1c815d038a4ed30a7d15d3bc9d5a177b8473e9417256
eval/bakeoff/copperline-ii/story_cards.json · CF4 Copperline Transit, year two · 200 cards
deb76619779aae48dc34e6f148591214ce4558991a1079cfeb1e9610b29882fd
eval/bakeoff/brinepool/story_cards.json · CF5 Brinepool · 320 cards

The round 4 registration · eval/bakeoff/ROUND4_REGISTRATION.sha256

Registered 7 September 2026 at 15:21:54 UTC, after the corpora were sealed and before the quiz was written. It names the hypothesis, the five contestants, the seven question classes, the three key policies the quiz was written to, the pins the run answers at, and the hash of the sealed quiz and key.

055432294cc2784775821828cc8150f92d542485944869fc91f15fcf03c31f94
eval/bakeoff/ROUND4_REGISTRATION.json · round 4's rules, written before the answers

How a reader verifies all four, in four commands

# from the repository root
shasum -a 256 -c eval/bakeoff/CORPORA_SEAL.sha256
shasum -a 256 -c eval/bakeoff/ROUND4_SEAL.sha256

# the registration names its file by basename,
# so it is checked from the folder it sits in
cd eval/bakeoff && shasum -a 256 -c WEEK0_REGISTRATION.sha256

# round 4's registration names its file from the repository root
cd .. && shasum -a 256 -c eval/bakeoff/ROUND4_REGISTRATION.sha256

On Linux the command is sha256sum -c and the arguments are the same. Fourteen lines come back, each ending OK. One line that does not is a broken seal, and a broken seal invalidates every number below it.

What a checksum proves, and what it does not

proves
The bytes of the corpus and of the registration are the bytes the seal was written against. Nothing has been edited to fit a result.
does not prove
When the seal was written. A hash file carries no clock and can be regenerated at any time.
what carries the date
The repository's history. The four corpus files, the registration and both checksum files entered it in one commit on 27 July 2026, and no commit since has touched any of them. git log --oneline -- eval/bakeoff/papers/story_cards.json shows one line.
Why the order matters: the registration was written first, the corpora were sealed half an hour later, and only then was the quiz written, by a separate agent that was told not to talk to Kit. The builder of the contestants never saw the questions or the answer keys. That order is the difference between a study and a demonstration, and it is the one thing about this bake-off that cannot be added afterwards.
3

The harness · five contestants, two metrics, one command anybody can run

Five readers answer the same questions over the same sealed cards under the same model, the same budget and the same instructions. Only the memory system behind the reader changes. Everything the run computes is derived from the rows it recorded, and a second program re-derives it from the file afterwards, so a figure in a published artifact is checkable by whoever holds the artifact.

The contestants

A no memory. B plain hybrid retrieval over the same cards, which the registration calls the honest baseline. C Kit retrieval only. D full Kit, with supersession and cite-or-abstain. E Graphiti, the named graph rival.

readers/contestants.py

The two metrics

Primary is claims: required claims present, forbidden claims absent, superseded guidance not repeated, and an abstention where the corpus genuinely has no answer. Secondary is cites, which also requires a resolvable source id. Both are always reported.

honesty.py · seven labels

Abstention is a result

A system that says it does not know scores on that question. A system that invents an answer loses on it. Seven labels separate a correct abstention from a missed one and an invention from a stale answer, so honesty is graded rather than assumed.

supported_correct, invent, stale, scope_leak, abstain_correct, abstain_miss, unsupported

What a query cost

Every answer carries its model, its tokens and its price, from one rate table and one rule. The reported figure for the Kit contestants is what the ledger charged, read back, not a second estimate of it.

readers/cost.py

What the corpus cost to load

Dollars per card on the way in, measured on both sides or on neither. One side metered is worse than none, because the asymmetry is the finding.

readers/write_cost.py

A grader that is not Kit

An optional second pass where a different frontier model sees the question, the rubric and one answer, blind to which contestant produced it, and replies pass or fail. It corroborates; it never replaces. Disagreements are listed, not resolved.

judge.py

One prompt for everybody

The answering instructions were a hidden variable: three copies in the tree, and the copy the Kit readers answered under carried clauses the baseline never saw. There is now one copy and a named profile, and a neutral rerun writes its own artifact.

readers/answer_prompt.py

How much is the sample

Every rate carries a 95% Wilson interval, and every gap against the baseline is tested paired on the questions both answered, exactly rather than through an approximation that wants far more disagreements than this quiz produces.

margin.py

How much is the asking

The whole quiz can be asked N times over, reporting the range beside the mean and naming the questions whose verdict moved. The temperature comes out of the seal rather than out of the code, and every row records the one it answered at.

repeats.py · readers/temperature.py

Reading an artifact back

Rows are the evidence and every other figure is derived, so the derived figures are recomputed from the rows and compared. A run audits the file it has just written and exits non-zero when the two disagree.

report.py

A read-back for someone with access to the study files

No database, no container, no answer key, no network, no API key. It reads one file in this repository and rebuilds every figure in it from the rows underneath. The artifact below is the apparatus run, which is Kit-authored and deliberately not the sealed study, and it can be checked by someone with access to those private repository files. It is not downloadable from this page.

python3 eval/bakeoff/report.py eval/bakeoff/mechanism/last_differentiated_run.json

B Plain RAG: claims 15/20   |  95% CI 53.1% to 88.8%
D Full Kit:  claims 18/20   |  95% CI 69.9% to 97.2%

D Full Kit: +15.0 points over 20 questions | won 3, lost 0, tied 17 | exact p 0.250
    clears the registered 10 point bar, and the paired test cannot
    tell this gap from chance
    about 37 questions at this rate would settle it (the study registered 120)

That last line is the harness reporting against its own product, unprompted, and it is the reason this page exists. --check prints the verdict alone and sets the exit code, so the check is a command rather than a reading exercise.

4

The numbers as they stand · round 4

Round 4 of the custodian's quiz: 120 questions in seven classes over the 1,246 sealed cards, answered under the tuned profile at the registered temperature, eight notes for every contestant, by Kit as it ships on 8 September (main at 000fafa5). The column to quote is the ruled one: the strict key's verdict, corrected only on the rows where the key and the blind judge disagreed and only by the five written policies, with 94 rows ruled and 1 left open. The key and the judge stand beside it, and the interval sits beside the rate rather than underneath it.

contestantruledrate95% intervalkeyjudgefound the noteread costwrite cost
C · Kit retrieval60 / 12050.0%41.2% to 58.8%366199 $1.11 per run $0.0018 per card
D · full Kit61 / 12050.8%42.0% to 59.6%406199 $1.12 per run $0.0018 per card
B · plain hybrid retrieval56 / 12046.7%38.0% to 55.6%385495 $1.05 per run a local embedding, no provider call
E · Graphiti43 / 12035.8%27.8% to 44.7%374248 $1.44 per run $0.051 per card at list
A · no memory30 / 12025.0%18.1% to 33.4%303230 nothing, by construction nothing, by construction

Found the note: how often the right card was among the eight shown, out of 120. Read cost is what the ledger charged for one pass over the quiz, answering and marking excluded. Write cost is what putting the corpus in cost per card, measured on both sides: Kit $2.19 for 1,246 cards, all of it the conflict judge; Graphiti $58.11 at list price for the 960 cards its graphs lacked, over 3 attempts under a $100 cap. Per card the rival's write path is about 29 times Kit's at list price, 34 times by the rate table.

The same quiz under the neutral profile, a plain answering prompt no contestant was coached under. The honest headline is this table beside the one above, not whichever flatters. 131 rows ruled, 1 left open.

contestantruledrate95% intervalkeyjudgefound the note
C · Kit retrieval67 / 12055.8%46.9% to 64.4%386499
D · full Kit65 / 12054.2%45.3% to 62.8%386099
B · plain hybrid retrieval59 / 12049.2%40.4% to 58.0%215795
E · Graphiti40 / 12033.3%25.5% to 42.2%364248
A · no memory30 / 12025.0%18.1% to 33.4%303030

What the margin is worth

the gap
+3.3 points for Kit retrieval over plain retrieval on the ruled column, tuned. Paired on the 120 questions: won 12, lost 8, tied 100, exact p 0.503. The registered pass bar is ten points, so on this column the gap does not clear it, and the test cannot tell it from chance. By the judge the gap is +5.8 points, which does not either.
the earlier resolver
Round 4f, the resolver that merged more, scored 64 and 59 under the rulings against 60 and 61 here. Paired, Kit retrieval -3.3 points from 4f to the shipped code (won 10, lost 14, exact p 0.541) and full Kit +1.7 (won 7, lost 5, p 0.774): inside the noise. What 4f's resolver cost is not on this exam: it retired 325 cards into merged notes and sent thirty conflicts to the operator; the shipped one retires none and sends none.
full Kit
+4.2 points over plain retrieval. Won 11, lost 6, tied 103, exact p 0.332. The gates that make full Kit cautious cost it answers on this exam; that thread is open.
the rival
-10.8 points for Graphiti against plain retrieval, and 14.2 points behind Kit retrieval. Kit retrieval against Graphiti: won 29, lost 12, tied 79, exact p 0.012. That gap is settled at this sample.
neutral
+6.7 points for Kit retrieval over plain retrieval on the ruled column under the neutral prompt. Won 12, lost 4, tied 104, exact p 0.077.
the sample
120 questions, the number the registration asked for. Each interval above is about eighteen points wide. Thirty of the 120 are questions whose right answer is a refusal, and the rival's score is mostly those.

What is measured, and what is only instrumented

measured
The grades under both profiles, the blind judge on every answer, the rulings on every disagreement but one, the read cost of every answer, and the write cost on both sides.
the rulings
Ruled by the five policies of 7 September, with the card in hand, by Kit's reader agents, one ruler and one skeptic per contestant. A row the skeptic would reverse is left open for the custodian rather than decided between two agents. The rulings file names every row, its policy and its verdict.
the rival's cost
Measured, at last. The graph rival read the round 4 cards with its own default models, one graph per case file, metered per extraction call, under a cap the custodian set. The record names every attempt, including the one that stopped when the provider account ran dry.
instrumented
Repeats. Asking the whole quiz several times over is code with tests and has not produced a number for round 4.

How the number got here, one fix at a time

Round 4 was built because rounds 1 to 3 could not separate Kit from a search box: 286 tidy cards, and nearly every question answerable from one card that shared its words. The new exam is messy on purpose, and Kit as it was found fewer of its own notes than plain retrieval did. Each fix below is one labelled run against the same sealed quiz, so a reader can see what each one did on its own. Key, judge and found, out of 120.

4
as Kit was, with a similarity floor of 0.35. Kit retrieval 38 by the key, 53 by the judge, 77 found; full Kit 34, 53, 77.
4b
the floor removed; the baseline never had one. Kit retrieval 37 by the key, 61 by the judge, 89 found; full Kit 38, 62, 89.
4c
keyword leg loosened to any word; worse, reverted. Kit retrieval 38 by the key, 53 by the judge, 86 found; full Kit 38, 54, 86.
4d
keyword leg weighted by word rarity; kept. Kit retrieval 39 by the key, 59 by the judge, 96 found; full Kit 36, 59, 96.
4e
conflict detector widened; conflicts resolved by the resolver as it was. Kit retrieval 38 by the key, 56 by the judge, 88 found; full Kit 39, 56, 88.
4f
the resolver keeps case, card and source tags on a merge and orders by event date. Kit retrieval 40 by the key, 67 by the judge, 98 found; full Kit 38, 64, 98.
4g
main as shipped on 8 September, the other builder's resolver: fewer merges, nothing forgotten, nothing escalated. Kit retrieval 36 by the key, 61 by the judge, 99 found; full Kit 40, 61, 99.

The lesson of 4e is the one worth keeping: a wider conflict net made Kit worse, because the resolver that merges two disagreeing notes was dropping their labels, so a merged note could not be found under its case file and the notes folded into it were gone. 4f fixed that and merged freely; 4g, the code that ships, merges only where it is sure, holds the rest open as uncertainty, forgets nothing and pages nobody, and gives back a few answers on the questions that need two months' notes read together. The key moves less than the judge across the series because it grades literal phrases and a merged note restates its sources; that is recorded as the key's problem, not the answer's.

Two things this table deliberately does not do. It does not quote the better of two profiles: the neutral table sits beside the tuned one and the reader takes both. And it does not quote the better of three graders: the ruled column is the key's verdict corrected only where the key and the judge disagreed and only by a written policy, and the key's own count stays in the row beside it.
5

The quiz stays sealed, and this is what that costs

The custodian's questions and answer keys are not published, and there is no plan to publish them. That decision was taken deliberately, with the cost understood, and stating the cost is part of taking it honestly.

Why

fitting
A published quiz is a quiz that can be fitted to. Every number measured over it afterwards is a number over a training target, and nobody outside can tell the difference.
the honesty labels
Are the easiest part to fit. The questions that expect an abstention are the ones a system most benefits from recognising in advance, and recognising a question is not remembering an answer.
the shelf life
A sealed quiz still means something in a year. A published one means something until the next model is trained.

What it costs, said plainly

the price
Nobody outside can reproduce the headline number. Not approximately, not with effort. That is the price of the seal and it is not softened here.
not the harness
The method is described here. The corpora, registration, readers, graders, cost model and statistics for this study live in Kit’s private repository. This page does not supply a public checkout. The separate 17-question demo exam is public and runnable.
not the failures
The disqualifier list is published with the study. Hiding failures, costs or abstentions ends the study rather than embarrassing it, and that sentence was written before there was anything to hide.
What an outsider does instead, which is the point: bring your own corpus and your own quiz. Seal them the same way, run the harness over them, and publish your seal beside your numbers. That is a harder thing to arrange than rerunning somebody else's quiz and a much better piece of evidence, because it tests the system rather than the exam. The number Kit intends to quote outward is the one an outsider produces that way, not the one in the table above.
The shape of the whole thing: the corpus is frozen and the freeze is checkable in one command · the rules were written before the answers · the builder never saw the quiz · every rate carries its interval and every gap carries its test · the bill is reported on both paths or on neither · what has not been measured says so · and the quiz stays sealed so that the result still means something after this page stops being new.