Skip to content
BetaBenchAlert v0.1 is in Beta.Numbers are real, but pages and rules can still change.See what changed

Claude Code

Haiku 4.5

Inside Claude Code we ask for claude-haiku-4-5 every hour. Everything below comes from those runs on our lab machine.

Claude Codeclaude-haiku-4-5Think setting highLast round 3m agoOther models

Write speed, last 24 hours

30.1answer tok/s

8.12s typical total wait

Likely middle: 6.62s9.03s

Visible answer tokens per second across the whole command. Hidden thinking does not make this number larger. Higher is faster. These are the middle values of 22 completed checks out of 24 scheduled in the last 24 hours.

Time to first word
4.52s
How long before the first visible word shows up. Lower is better.
Typical total wait
8.12s
How long the whole command took, start to finish, including the app booting. Lower is faster.
Last ping
3.21s
The newest tiny “reply ok” check. It measures wake-up time, not writing speed.

What stands out

Four facts pulled straight from the saved runs.

Board position
#6 of 10
Just ahead: Fable 5 at 7.79s total wait. Just behind: GPT-5.6 Terra at 10.4s total wait.
Answered when scheduled
22/24
Completed answers out of all scheduled checks in the last 24 hours.
Served as
claude-haiku-4-5
The model name the app reported back. If it differs from the pin, we say so.
Round-to-round swing
1.8×
How much the fastest round beat the slowest one. A big swing means a noisy day.

Recent rounds

Newest first. Fails stay in the list with their reason.

WhenResultWriting speedTotal wait
3m agookclean40.8 tok/s6.00s
3h agookclean33.8 tok/s7.22s
4h agookclean38.4 tok/s6.39s
5h agookclean32.5 tok/s7.55s
6h agookclean30.4 tok/s8.02s
7h agookclean43.4 tok/s5.62s
8h agookclean38.4 tok/s6.36s
9h agookclean30.2 tok/s8.11s

All rounds are in Logs.

Read next

Compare with GPT-5.6 Luna

Its closest rival in Codex — same window, same units, with the leader marked on every row.

Compare with Sonnet 5

The pin next to it in Claude Code. See what the tier jump buys inside one app.

More models

How to read this

Every number here comes from our lab machine running the paid app, one try per hourly round. The live check uses test kit v0.2; recovered history appears only after its raw file passes the same visible-speed rules. Our measurements, not the vendor’s.