skip to note
back to notes

apr 18–19, 2026 · atlas

atlas beat baseline

28 tasks × 5 trials · +0.125 det · −34% tokens/task

atlas benchmark table comparing four code-research agents
atlas benchmark table comparing four code-research agents
timestamp
at 18:03 on apr 18, 58 seconds after the benchmark commit, the score was +0.125 deterministic and +0.035 llm while tokens per task fell 34% and cost fell 25%.
benchmark setup
i mapped symbols, calls, and tests, then compared the agent against plain text tools before trusting the graph.
suspicious result
the first result looked too good, so i gave the judge code tools and added unseen corpora before trusting it.
the pattern
i turn an idea into a tool, measure it, look for the embarrassing counterexample, and keep whatever survives.

tags

  • #code intelligence
  • #evals
  • #sqlite graph
  • #mcp