Leaderboard

Usage Units per configuration.

Compare the harness, model, and effort level independently—then inspect validated work, task time, completion, and throughput without exposing benchmark internals.

Rank by
Configurations
—
Harness × model × effort
Usage Units
—
Validated work published
Attempts
—
Externally checked tasks
Task time
—
Measured task wall time

Comparison graphs

Weekly-equivalent work
WEEKLY EQUIVALENT UU · HIGHER IS BETTER
Work versus task time

Higher and farther left is better.

Configuration telemetry

Only safe run-level aggregates are public; unavailable provider telemetry is never inferred.

#HarnessModelEffortLimitPlanUUWeekly equivalent CompletionAttemptsTask timeUU/hour