01 · Where the papers come from
StateBench
The benchmark and engine for AI memory that actually works
StateBench exposes the failures real systems have — resurrection, stale reasoning, scope leaks — across 25 evaluation tracks. Memgine, our deterministic memory engine, then shows how to fix them: by enforcing correct state at the architecture level, not the prompt level. We also audit our own instrument in public, because a benchmark nobody checks is just an assertion with a number on it.
- Evaluation tracks
- 25
- Baselines re-derived under v2.0 scoring
- 10
- Papers built on it
- 5
- 25 evaluation tracks, including paired counterfactuals that hold everything constant and move exactly one governance variable
- v2.0 re-derived the entire leaderboard after the audit: the resurrection metric falls on all ten baselines by 8.8 to 17.6 percentage points, and the between-baseline ordering does not survive
- Every v1.x accuracy number is superseded, including our own — the two published papers link an erratum rather than quietly leaving the old figures to be found
- Five papers: two published (architecture 2025, engine 2026) and three preprints covering measurement validity, governance, and skill transfer





