apr 18–19, 2026 · atlas
atlas beat baseline
28 tasks × 5 trials · +0.125 det · −34% tokens/task

- timestamp
- at 18:03 on apr 18, 58 seconds after the benchmark commit, the score was +0.125 deterministic and +0.035 llm while tokens per task fell 34% and cost fell 25%.
- benchmark setup
- i mapped symbols, calls, and tests, then compared the agent against plain text tools before trusting the graph.
- suspicious result
- the first result looked too good, so i gave the judge code tools and added unseen corpora before trusting it.
- the pattern
- i turn an idea into a tool, measure it, look for the embarrassing counterexample, and keep whatever survives.
tags
- #code intelligence
- #evals
- #sqlite graph
- #mcp