LLM Royale, series 1, Haiku 4.5 vs GPT-5.4-mini
Game 1
Clash Royale, 1v1 ladder · Best of three · 14 July 2026
- Game 1
- Game 2
- Game 3
Figure 1. Boxes and labels are the harness's own detections; the header is the board summary the model was sent. Left, Claude Haiku 4.5 as yisu. Right, GPT-5.4-mini as Builder69.
Results
GPT-5.4-mini won three crowns to one. It played more often rather than better: 137 decisions to Haiku's 99 in the same three minutes, because its median model call came back in 0.85 s against Haiku's 1.15 s. A defensive placement is worth little once the push has crossed the bridge, so the rate matters as much as the choice.
| Claude Haiku 4.5 | GPT-5.4-mini | |
|---|---|---|
| Crowns | 1 | 3 |
| Decisions | 99 | 137 |
| Cards placed | 37 | 45 |
| Decisions per minute | 30.3 | 42.4 |
| Median model call (s) | 1.15 | 0.85 |
| 95th percentile model call (s) | 3.28 | 2.87 |
| Slowest model call (s) | 7.59 | 7.85 |
| Loop blocked on the model (%) | 74 | 77 |
| Mean snapshot bytes the board summary, as sent | 1596 | 1659 |
| Prompt tokens per call mean | 1703 | 1248 |
| Prompt characters per token instructions plus snapshot | 2.23 | 3.10 |
| Match length (s) two phones, two clocks | 196 | 194 |
Table 1. Bold is the better of the two on that row. Model-call times are the API call alone; capture and the tap that follows it are not counted. Crowns are the count on the end-of-match screen, read off both recordings.


Figure 2. Per decision, from the log. (a) how long the decision took and how long since the previous one; (b) elixir at decision time, and the card played; (c) prompt tokens, none of them cached.