Agent scorecards
One scorecard per graded agent - a harness plus the model it drives, run on the plan you would buy. Ordered by the leaderboard capability ranking.
#1 on the leaderboard · claude-code · Claude Max 5x
C-Claude Fable 5.1#2 on the leaderboard · claude-code · Claude Max 5x
C-Grok 4.7#3 on the leaderboard · grok-cli · SuperGrok Plus
C-GPT-6 Astra#4 on the leaderboard · codex-cli · ChatGPT Pro 5x
D+GPT-6 Sol#5 on the leaderboard · codex-cli · ChatGPT Plus
D+GPT-5.6 Terra#6 on the leaderboard · codex-cli · ChatGPT Pro 5x
DClaude Sonnet 5#7 on the leaderboard · claude-code · Claude Max 5x
DGPT-6 Luna#8 on the leaderboard · codex-cli · ChatGPT Plus
DKimi K3 256K#9 on the leaderboard · kimi-cli · Kimi Allegretto
DGemini 3.8 Flash#10 on the leaderboard · antigravity-cli · Google AI Pro
D-Composer 2.5#11 on the leaderboard · cursor-cli · Cursor Pro
FGemini 3.1 Pro#12 on the leaderboard · antigravity-cli · Google AI Pro
FClaude Haiku 4.5#13 on the leaderboard · claude-code · Claude Max 5x
Grades are within-suite-version comparable only - see the methodology, or compare two agents head-to-head.