Independent capacity measurement

How much work does a plan really do?

UsageBenchmark measures how much real, validated coding work an AI coding plan completes before it hits its usage limit — and reports it as a single number: the Usage Unit. Not tokens. Not requests. Not marketing.

Work you can count, not usage you can't.

1 UU

One Usage Unit is one standardized coding task, completed and confirmed by an independent check. Bigger, harder tasks are worth more units. A plan's score is the total it finishes before its limit.

Providers describe their plans in tokens, messages, or vague "usage." Those numbers don't tell you what you actually get done. UsageBenchmark answers the practical question instead: run a real coding agent on a fixed set of tasks until the plan taps out, and count the work that genuinely passed.

01

Real coding work

The agent solves genuine programming tasks across several languages — the kind of work these plans are actually bought for.

02

Independently scored

Every task is judged by an automatic external check. The agent cannot mark its own work as done; only confirmed work earns units.

03

Run to the limit

Each plan runs until it reaches its own usage limit. The published score is the validated work completed within that window.

Only the results are public.

This is a private benchmark run by one operator. The tasks, the scoring checks, and the exact methodology are deliberately not published — if they were, a model could be tuned to the test and the numbers would stop meaning anything.

✕Not published
✓Published
✕The task set, prompts, and answer checks
✓Usage Units earned, per model and plan
✕How difficulty and units are calibrated
✓Completion rate and run date
✕The runner, schedule, and internal evidence
✓Whether the run hit the plan's usage limit

The result is a comparison you can trust precisely because the recipe stays sealed: the published Usage Units reflect what a plan did on tasks it had never seen.