Real tests of the AI coding tools you pay for.
Claude Code, Codex, and Grok Build — timed on the same paid plans you buy. How fast they write, and whether the answer is right. For people on a subscription, not an API key.
Livejust now
- Fable 5overloadedlast @ 12:00 AMhigh effort54.85 tok/s4.39s total
- Opus 5overloadedlast @ 1:00 AMhigh effort41.01 tok/s5.88s total
- Grok 4.6xhigh effort40.97 tok/s5.83s total
- GPT-5.6 Solhigh effort24.58 tok/s9.89s total
- Measured
- 2:04 AM
- Timezone
- ET
- Test kit
- v0.2
—
Quota
What 100% of a weekly limit is worth at the vendor's own API prices.
Liveupdated 2h agonext run 03:10 UTC
- Claude Max 20xdegraded≈ $2,636
5-hour ≈ $575 · Fable 5 up to 50% of weekly
- SuperGrok Heavydegraded≈ $1,232
Includes Cursor Ultra — $400/mo usage on top
- ChatGPT Pro ($200, 20×)no reading yet—
- Last reading
- Aug 24
- Method
- v0.1
- Interval
- daily
Highlights
The last 24 hours, made simple
Live speed is in the hero. Here we step back and ask two steadier questions: what was typical, and how often did each scheduled check return an answer?
Typical writing speed
Visible answer tokens per second, last 24 h · higher is faster
4 of 10 models
Other modelsAnswered when scheduled
Completed answers out of all scheduled checks, last 24 h · higher is better
4 of 10 models
Lab statusBoth boards cover the last 24 hours. Each board has its own scale · test kit v0.2. 4 of 10 pinned models shown · the rest live on Other models.
Speed
Typical speed and total wait
The 4 models we lead with, ranked by typical total wait over the last 24 hours. At least 12 completed checks are required.
Writing speed counts only the answer you can see. Total wait measures click-to-finish time. The table ranks the shorter total wait first.
| Rank | Model | Bar, on one shared scale | Writing speed | Total wait | vs yesterday | Answered | Coverage |
|---|---|---|---|---|---|---|---|
| 1 | Grok 4.6xAI · effort xhigh | 36.55 | 6.54slikely 6.11s–8.08s | -2% | 100% | 24/24 | |
| 2 | Opus 5Anthropic · effort high · last round overloaded | 34.74 | 6.94slikely 5.32s–7.35s | +1% | 96% | 24/24 | |
| 3 | Fable 5Anthropic · effort high · last round overloaded | 30.80 | 7.83slikely 5.59s–8.17s | -8% | 92% | 23/24 | |
| 4 | GPT-5.6 SolOpenAI · effort high | 22.23 | 10.9slikely 10.5s–11.8s | +1% | 88% | 24/24 |
Window 24 hours · 4 of 4 models have 12 completed checks · 95 checks attempted · test kit v0.2. A dash under vs yesterday means that model did not have 12 completed checks the day before. 4 of 10 pinned models shown · the rest live on Other models.
24 hours
The last 24 hours, round by round
Each line joins one model's rounds, one an hour. Every dot is a real round, so the swings stay in plain sight. Point at the plot to read any round.
Visible answer tokens per second · higher is faster
Model writing speed over time
Lab time (ET)
Line: one model's visible writing speed, check to check. Dots: the checks themselves. A single missed hour is stepped over by a faint dotted link; anything longer breaks the line. Nothing is invented to fill a gap.
Full record — every round, every number
| Round | Grok 4.6 | Opus 5 | Fable 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Aug 24 2:00am | 40.97 | — | — | 24.58 |
| Aug 24 1:00am | 21.62 | 41.01 | — | 24.90 |
| Aug 24 12:00am | 38.22 | 45.42 | 54.85 | 23.20 |
| Aug 23 11:00pm | 41.64 | 51.93 | 41.69 | 25.30 |
| Aug 23 10:00pm | 40.16 | 34.74 | 49.94 | 22.53 |
| Aug 23 9:00pm | 28.10 | 57.39 | 43.09 | 21.07 |
| Aug 23 8:00pm | 14.35 | 49.77 | 48.18 | 20.68 |
| Aug 23 7:00pm | 33.64 | 45.33 | 53.26 | 23.27 |
| Aug 23 6:00pm | 40.88 | 52.93 | 43.33 | — |
| Aug 23 5:00pm | 32.96 | 50.25 | 37.34 | — |
| Aug 23 4:00pm | 36.02 | 35.21 | 28.60 | — |
| Aug 23 3:00pm | 39.15 | 33.24 | 26.33 | 19.60 |
| Aug 23 2:00pm | 28.01 | 33.30 | 30.53 | 22.23 |
| Aug 23 1:00pm | 40.08 | 34.58 | 31.07 | 15.38 |
| Aug 23 12:00pm | 29.57 | 33.47 | 24.05 | 20.31 |
| Aug 23 11:00am | 39.17 | 29.45 | 32.81 | 23.15 |
| Aug 23 10:00am | 37.17 | 35.56 | 26.18 | 22.92 |
| Aug 23 9:00am | 14.83 | 30.96 | 30.00 | 20.17 |
| Aug 23 8:00am | 28.73 | 31.05 | 30.49 | 19.51 |
| Aug 23 7:00am | 37.83 | 28.55 | 31.61 | 20.57 |
| Aug 23 6:00am | 34.69 | 38.97 | 30.21 | 23.01 |
| Aug 23 5:00am | 37.01 | 31.80 | 29.24 | 21.88 |
| Aug 23 4:00am | 37.65 | 32.77 | 29.49 | 21.18 |
| Aug 23 3:00am | 36.09 | 28.15 | 30.51 | 24.59 |
Window 24 hours · one round every 60 minutes · 24 rounds on the clock · n = 90 completed readings drawn · test kit v0.2. 4 of 10 pinned models shown · the rest live on Other models.
We also ping each app every round to check it answers at all. Those checks, round by round, live on Lab status.
Method
How we get these numbers
We run the real paid apps on our own lab machine, not the hidden APIs. Every number traces back to a saved run.
- Cadence
- Every hour
- Rounds a day
- 24
- Test kit
- v0.2
- Window
- 24hours
Last round 5m ago · 309 results saved in the last 24 hours, 27 failed · Lab status.