xAI released Grok 4.7 on September 21. It calls it its most powerful model for coding and knowledge work, "twice as fast, at half the price of comparable models", served at the same price and speed as Grok 4.6: $2 per million input tokens and $6 per million output, half of Opus 5.5's input price and under a third of its output price. We graded it as the successor to Grok 4.6.
What the results show
- It reads as a tier switch, not a new model. We grade Grok at high reasoning effort, the setting 4.6 ran at. At that setting 4.7 scored better on real issues than 4.6 had, but only by working far longer: tasks ran into our time limit on real issues as well as on planning, and its grading run took three to four times as long as the next slowest agent's. We had to lower it to medium effort to get through the suites at all. At medium, Grok 4.7 lands exactly where 4.6 was: the same rung on real issues, the same plan-writing grade, and one rung more on vibe-coding, where it traded a task here and there with 4.6. What xAI calls 4.7 at medium behaves like 4.6 at high.
- On grade alone it sits third. Behind Claude Opus 5.5 and Claude Fable 5.1, level with GPT-6 Astra. That is the grade, not the experience of using it.
- It is light on the plan. A run of our suites takes about 12% of the SuperGrok Plus weekly limit, about what 4.6 used. At the same monthly price a run takes about 6% of Astra's weekly limit.
The wait is the story. Even at medium effort a real issue takes Grok 4.7 about seven minutes, the same as 4.6 and about three times as long as Opus 5.5 or Astra, which makes it the slowest agent on the bench by a wide margin. At high effort the median was twelve minutes per issue, real issues and planning tasks alike ran out the clock at forty minutes, and the whole grading run took three to four times as long as the next slowest agent's. No other provider has ever hit that limit in our runs. Whatever the extra thinking buys on a benchmark, nobody waits twelve minutes for a fix to a one-line bug.
Our take: Grok 4.5 was fast enough to be useful for everyday work. 4.6 gave that up and 4.7 goes further the same way, with the effort dial doing the work a new model should have done. Too slow for routine work and not capable enough for difficult work, so we see no reason to use it. The $30 and $100 SuperGrok plans buy chat, images and video, not a coding agent we would use.
Grades are comparable within the same suite version. See the full Grok 4.7 scorecard or follow the family over time on Trends.