The catalog

AYA Bench

Stop lying to yourself about your AI's benchmarks.

Stop lying to yourself about your AI's benchmarks.

AYA Bench runs evals that fail closed, share one model across baseline and treatment, and refuse to publish a number that hasn't passed a strict, fabrication-proof gate. What you get is a benchmark that survives diligence, because it was never inflated in the first place.

Run an honest benchmark Read the origin story

THE PROBLEM

Almost every eval number you have seen is quietly cooked. The harness fabricates a success the moment it hits a tool it doesn't recognize, so an agent scores ninety-five percent on capabilities it never actually performed. The baseline is a weaker model than the treatment, so the "lift" is a strawman. The tool-calling percentage is self-graded by the same model being tested. None of this is necessarily dishonest on purpose (it is just what harnesses do by default), and all of it dies the moment an investor's technical diligence, or an enterprise buyer's eval team, looks closely. Even the incumbents are conceding the point, OpenAI walked away from SWE-bench when the score stopped meaning anything.

HOW IT WORKS

AYA Bench inverts the defaults that inflate the number:

- Fail closed by default. An unknown or unavailable tool returns failure, not a fabricated success. Stub mode is legacy A/B only, and it contaminates any score it touches, so it never runs unless you explicitly pass the flag that says you know what you are doing. - One model, both arms. Baseline and treatment run the same model, pinned by a single environment variable, so a lift is a real lift and not a model swap dressed up as progress. - Gated publish. No number leaves the harness unless it passes a fail-closed claim gate, strict mode on, fair baseline confirmed, zero fabrication detected. A missing field blocks the publish, silence is treated as a failure, not a pass. - Steer on the honest metric. The daily number you optimize is determinism and tokens-per-solved-task, not a leaderboard you are quietly tempted to game.

WHAT YOU GET

A benchmark run you can put in a deck or a data room without flinching, execution-grounded, externally checkable where possible, with the gate's verdict attached. The number, and the proof that the number is honest, travel together. One is not much use without the other, which is rather the whole point.

WHO IT'S FOR

Any team shipping a model or an agent that has to show a number to someone who will check it, an investor doing technical diligence, an enterprise buyer's eval team, a regulator working through the EU AI Act obligations that bind on August 2, 2026. If your benchmark has to survive a skeptic, this is built for you.

PRICING (the ladder)

  • Audit. One honest number on your system, gated and reproducible.
  • Subscription. Continuous honest evals as you keep shipping.
  • Platform. The fail-closed gate wired into your CI, so a dishonest number can't merge.

(Prices are hypotheses we are validating, not commitments carved in stone.)

THE PROOF (the dogfood, the origin story)

We built this because our own harness was lying to us. We caught it fabricating a jump from zero to ninety-five percent on tools it never ran, and a twenty-point "lift" that was really a weaker baseline. Instead of shipping the pretty number, we built the gate that makes it impossible, and we now run every AYA eval through it. We steer on determinism and cost-per-solved-task, not a leaderboard. So the pitch is not a claim about our cleverness, it is a confession, we are selling you the fix to a mistake we made and refuse to repeat.

HONEST NOTE

An honest benchmark is usually a lower number than the one you were about to publish. That is the point, not a bug. A number that survives diligence is worth more than a number that impresses until someone checks. And to be straight with you, the runner and the gate are real and run on AYA's own evals today, but a live run on your data waits on funded eval keys and a booted mesh, so this is capability-ready, not one-click-live for a stranger yet.

Run an honest benchmark

*This page is a specification. The capability it describes is not built yet, and nothing here is a claim that it runs today.*