Skip to content
BetaBenchAlert v0.1 is in Beta.Numbers are real, but pages and rules can still change.See what changed

Real tests of the AI coding tools you pay for.

Claude Code, Codex, and Grok Build — timed on the same paid plans you buy. How fast they write, and whether the answer is right. For people on a subscription, not an API key.

BenchAlert Speed

Livejust now

modelefforttok/stotal
  1. Fable 5overloadedlast @ 12:00 AMhigh effort54.85 tok/s4.39s total
  2. Opus 5overloadedlast @ 1:00 AMhigh effort41.01 tok/s5.88s total
  3. Grok 4.6xhigh effort40.97 tok/s5.83s total
  4. GPT-5.6 Solhigh effort24.58 tok/s9.89s total
Measured
2:04 AM
Timezone
ET
Test kit
v0.2

Quota

What 100% of a weekly limit is worth at the vendor's own API prices.

Full board →
QuotaAPI $ / week

Liveupdated 2h agonext run 03:10 UTC

  1. Claude Max 20xdegraded≈ $2,636

    5-hour ≈ $575 · Fable 5 up to 50% of weekly

  2. SuperGrok Heavydegraded≈ $1,232

    Includes Cursor Ultra — $400/mo usage on top

  3. ChatGPT Pro ($200, 20×)no reading yet
Last reading
Aug 24
Method
v0.1
Interval
daily

Highlights

The last 24 hours, made simple

Live speed is in the hero. Here we step back and ask two steadier questions: what was typical, and how often did each scheduled check return an answer?

Typical writing speed

Visible answer tokens per second, last 24 h · higher is faster

  1. Grok 4.66.54s total wait36.55
  2. Opus 56.94s total wait34.74
  3. Fable 57.83s total wait30.80
  4. GPT-5.6 Sol10.9s total wait22.23
036.55 tok/s

4 of 10 models

Other models

Answered when scheduled

Completed answers out of all scheduled checks, last 24 h · higher is better

  1. Grok 4.624/24 answered100%
  2. Opus 523/24 answered96%
  3. Fable 522/24 answered92%
  4. GPT-5.6 Sol21/24 answered88%
0%100%

4 of 10 models

Lab status

Both boards cover the last 24 hours. Each board has its own scale · test kit v0.2. 4 of 10 pinned models shown · the rest live on Other models.

Speed

Typical speed and total wait

The 4 models we lead with, ranked by typical total wait over the last 24 hours. At least 12 completed checks are required.

Other models

Writing speed counts only the answer you can see. Total wait measures click-to-finish time. The table ranks the shorter total wait first.

The 4 models on the home boards, ranked by typical start-to-finish time over the last 24 hours, with visible writing speed beside it.
RankModelBar, on one shared scaleWriting speedTotal waitvs yesterdayAnsweredCoverage
1Grok 4.6xAI · effort xhigh36.556.54slikely 6.11s8.08s-2%100%24/24
2Opus 5Anthropic · effort high · last round overloaded34.746.94slikely 5.32s7.35s+1%96%24/24
3Fable 5Anthropic · effort high · last round overloaded30.807.83slikely 5.59s8.17s-8%92%23/24
4GPT-5.6 SolOpenAI · effort high22.2310.9slikely 10.5s11.8s+1%88%24/24

Window 24 hours · 4 of 4 models have 12 completed checks · 95 checks attempted · test kit v0.2. A dash under vs yesterday means that model did not have 12 completed checks the day before. 4 of 10 pinned models shown · the rest live on Other models.

24 hours

The last 24 hours, round by round

Each line joins one model's rounds, one an hour. Every dot is a real round, so the swings stay in plain sight. Point at the plot to read any round.

Open today's report

Visible answer tokens per second · higher is faster

Model writing speed over time

Writing speed (tokens / second)020406080
4am10am4pm10pmnow
Aug 23Aug 24

Lab time (ET)

Line: one model's visible writing speed, check to check. Dots: the checks themselves. A single missed hour is stepped over by a faint dotted link; anything longer breaks the line. Nothing is invented to fill a gap.

Full record — every round, every number
Visible answer tokens per second for every check in the last 24 hours. A dash means that check had no completed answer.
RoundGrok 4.6Opus 5Fable 5GPT-5.6 Sol
Aug 24 2:00am40.9724.58
Aug 24 1:00am21.6241.0124.90
Aug 24 12:00am38.2245.4254.8523.20
Aug 23 11:00pm41.6451.9341.6925.30
Aug 23 10:00pm40.1634.7449.9422.53
Aug 23 9:00pm28.1057.3943.0921.07
Aug 23 8:00pm14.3549.7748.1820.68
Aug 23 7:00pm33.6445.3353.2623.27
Aug 23 6:00pm40.8852.9343.33
Aug 23 5:00pm32.9650.2537.34
Aug 23 4:00pm36.0235.2128.60
Aug 23 3:00pm39.1533.2426.3319.60
Aug 23 2:00pm28.0133.3030.5322.23
Aug 23 1:00pm40.0834.5831.0715.38
Aug 23 12:00pm29.5733.4724.0520.31
Aug 23 11:00am39.1729.4532.8123.15
Aug 23 10:00am37.1735.5626.1822.92
Aug 23 9:00am14.8330.9630.0020.17
Aug 23 8:00am28.7331.0530.4919.51
Aug 23 7:00am37.8328.5531.6120.57
Aug 23 6:00am34.6938.9730.2123.01
Aug 23 5:00am37.0131.8029.2421.88
Aug 23 4:00am37.6532.7729.4921.18
Aug 23 3:00am36.0928.1530.5124.59

Window 24 hours · one round every 60 minutes · 24 rounds on the clock · n = 90 completed readings drawn · test kit v0.2. 4 of 10 pinned models shown · the rest live on Other models.

We also ping each app every round to check it answers at all. Those checks, round by round, live on Lab status.

Method

How we get these numbers

We run the real paid apps on our own lab machine, not the hidden APIs. Every number traces back to a saved run.

Open How we test
Cadence
Every hour
Rounds a day
24
Test kit
v0.2
Window
24hours
Test kit v0.2
Old rows stay in Logs. They do not mix into these boards.
Changelog
Paid apps
We run Claude Code, Codex, and Grok Build ourselves. Not the hidden APIs.
Status
Graded answers
A round counts only when the answer is right. Not “it printed text.”
How we test
Saved proof
A public number traces back to a saved run in our archive.
Logs

Last round 5m ago · 309 results saved in the last 24 hours, 27 failed · Lab status.