See what each model builds when the world keeps moving.
Explore the current gameplay recordings, then inspect historical benchmark results separately. Different interfaces and recording windows are not a model ranking.
Same start. Same controls. Six attempts.
Astra, Sol and Terra each get two fresh attempts from the same fortress seed. All use Medium reasoning, the same prompt and a 120 × 40 text screen. Compare equal decision budgets, not equal game time or token use. At 128, each model continues its own save and memory.
Loading saved experiment results…
Blank results mean no reviewed result is published here yet, not failure or zero progress. An unfinished run may retain an earlier checkpoint. Infrastructure failures are not model gameplay failures. Watch the active session →
Supplies are saved stock counts, not production rates or proof the dwarves can access them. Model charges are unreported by the subscription, not $0. These early, single-seed attempts do not establish a model ranking.
Read the frozen experiment plan → · Download the selected results
Keyboard controls or workshop shortcuts?
Three paired repeats from the same starting save, with empty starting memory and the same 120 × 40 screen. Both conditions allow standard keyboard input. The shortcut condition adds selected-workshop job insertion and its instructions. Each attempt has the same 128-decision budget.
6 recorded results · 3 pairs · 128 decisions each
| Pair / controls | Elapsed game ticks | Living / recorded deaths | Completed beds / workshops / farms | Raw food / drinks | Returned tokens | Evidence |
|---|---|---|---|---|---|---|
| Pair 1Keyboard | 70,900 | 7 / 0 | 7 / 2 / 1 | 43 / 135 | 3,879,799 | ReplayResult |
| Pair 1Keyboard + shortcuts | 68,200 | 7 / 0 | 7 / 2 / 1 | 39 / 115 | 3,772,495 | ReplayResult |
| Pair 2Keyboard | 51,900 | 6 / 1 | 7 / 5 / 1 | 61 / 122 | 3,954,012 | ReplayResult |
| Pair 2Keyboard + shortcuts | 49,600 | 7 / 0 | 10 / 2 / 2 | 40 / 136 | 4,014,250 | ReplayResult |
| Pair 3Keyboard | 46,400 | 7 / 0 | 7 / 3 / 1 | 42 / 122 | 3,753,509 | ReplayResult |
| Pair 3Keyboard + shortcuts | 58,200 | 7 / 0 | 7 / 2 / 1 | 56 / 161 | 3,621,402 | ReplayResult |
Neither condition clearly dominates these endpoints. These are repeated attempts on one world, not independent seeds or a strong controls ranking. Equal decisions do not mean equal game time or token use. These short attempts are separate from the Year-Two endurance run.
A queued job is not a completed product. Saved supplies are not production rates, accessible reserves or proof of self-sufficiency. The six attempts used 22,995,467 returned tokens; dollar charges are unreported, not $0.
Each attempt reloaded its own decision-64 checkpoint before continuing. The final decision-128 saves do not claim another independent reload.
Read the declared conditions → · Download all paired results · Reproduce the comparison →
Watch the play. Keep the comparison honest.
Replay captured screens, chosen keys and game responses. Older recordings use different control conditions and are separate from the matched attempts above. Astra’s earlier continuation reached Year Two; it is not a repeat in this new comparison.
Comparable results need a shared condition.
Shared worlds, prompts, memory, interfaces, models, code, and evaluator versions define one complete Easy v1 field.
Open the frozen G7-v3 protocol →The score archive keeps its original context.
Explore earlier score versions and seed lanes in the historical dashboard. They remain separate from Fort-Eval cohorts.
Every score tells a story you can replay.
Each replay here is selected directly from the protocol field above.