Electricity Bench

Claude Fable 5.1 beats out GPT-6 Astra - Week 37 agent results

Weekly Run

By the Electricity Bench editorial team · Drafted with AI, reviewed and approved by a human editor

The first weekly grading run after the release of Claude Fable 5.1, GPT-6 Astra and Gemini 3.8 Flash side by side.

What a crazy week behind us. We started off with the release of Claude Fable 5.1, then came Gemini 3.8 and finally GPT-6 Astra. The week 2026-W37 grading run is the first weekly run to grade them side by side, across real-world@v3, spec-planning@v1 and vibe-coding@v1.

What the results show

The week goes to Claude Fable 5.1, with GPT-6 Astra right behind. The part we keep thinking about is the cheap model. From experience, several models got worse after the release of a newer version. Our suspicion is that providers trim a model to save money once its successor ships, through heavier quantization or a lower reasoning effort. We cannot prove this, so read it as a guess. With that in mind we have to find a successor for Gemini 3.7 as our driver for smaller tasks (our take).

Grades are comparable within the same suite version. Explore the complete Week 37 results or follow each agent over time on Trends.