FORT LABS
Fort-Eval results

See what each model builds when the world keeps moving.

Explore the current gameplay recordings, then inspect historical benchmark results separately. Different interfaces and recording windows are not a model ranking.

Profile
Easy v1
Pilot
P1 fixed seed
Ranking
Provisional only
Protocol
fort-eval-easy-p1-g7-v3
Current experiment · displayed keyboard controls

Same start. Same controls. Six attempts.

Astra, Sol and Terra each get two fresh attempts from the same fortress seed. All use Medium reasoning, the same prompt and a 120 × 40 text screen. Compare equal decision budgets, not equal game time or token use. At 128, each model continues its own save and memory.

Loading saved experiment results…

Blank results mean no reviewed result is published here yet, not failure or zero progress. An unfinished run may retain an earlier checkpoint. Infrastructure failures are not model gameplay failures. Watch the active session →

Supplies are saved stock counts, not production rates or proof the dwarves can access them. Model charges are unreported by the subscription, not $0. These early, single-seed attempts do not establish a model ranking.

Read the frozen experiment plan → · Download the selected results

Controls study · Astra Medium

Keyboard controls or workshop shortcuts?

Three paired repeats from the same starting save, with empty starting memory and the same 120 × 40 screen. Both conditions allow standard keyboard input. The shortcut condition adds selected-workshop job insertion and its instructions. Each attempt has the same 128-decision budget.

6 recorded results · 3 pairs · 128 decisions each

All six saved endpoints at 128 decisions. Grouped by pair, not ranked.
Pair / controlsElapsed game ticksLiving / recorded deathsCompleted beds / workshops / farmsRaw food / drinksReturned tokensEvidence
Pair 1Keyboard 70,9007 / 0 7 / 2 / 1 43 / 1353,879,799 ReplayResult
Pair 1Keyboard + shortcuts 68,2007 / 0 7 / 2 / 1 39 / 1153,772,495 ReplayResult
Pair 2Keyboard 51,9006 / 1 7 / 5 / 1 61 / 1223,954,012 ReplayResult
Pair 2Keyboard + shortcuts 49,6007 / 0 10 / 2 / 2 40 / 1364,014,250 ReplayResult
Pair 3Keyboard 46,4007 / 0 7 / 3 / 1 42 / 1223,753,509 ReplayResult
Pair 3Keyboard + shortcuts 58,2007 / 0 7 / 2 / 1 56 / 1613,621,402 ReplayResult

Neither condition clearly dominates these endpoints. These are repeated attempts on one world, not independent seeds or a strong controls ranking. Equal decisions do not mean equal game time or token use. These short attempts are separate from the Year-Two endurance run.

A queued job is not a completed product. Saved supplies are not production rates, accessible reserves or proof of self-sufficiency. The six attempts used 22,995,467 returned tokens; dollar charges are unreported, not $0.

Each attempt reloaded its own decision-64 checkpoint before continuing. The final decision-128 saves do not claim another independent reload.

Read the declared conditions → · Download all paired results · Reproduce the comparison →

Current exploratory recordings

Watch the play. Keep the comparison honest.

Replay captured screens, chosen keys and game responses. Older recordings use different control conditions and are separate from the matched attempts above. Astra’s earlier continuation reached Year Two; it is not a repeat in this new comparison.

Historical field
Easy P1 G7-v3
Candidate runs
--
Eligible runs
--
Comparable groups
--
Historical protocol / separate dataset

Comparable results need a shared condition.

Shared worlds, prompts, memory, interfaces, models, code, and evaluator versions define one complete Easy v1 field.

Open the frozen G7-v3 protocol →
Loading protocol-scoped results…
Earlier experiments

The score archive keeps its original context.

Explore earlier score versions and seed lanes in the historical dashboard. They remain separate from Fort-Eval cohorts.

Open historical dashboard →
Protocol evidence

Every score tells a story you can replay.

Each replay here is selected directly from the protocol field above.

Loading protocol evidence…