Vol. I · No. 1 An Essay in Game Studies

The Ludic Mirror,
Revisited

A critique of an ambitious treatment, and a more honest blueprint for simulating worlds and the meanings they make

The original treatment is intellectually generous and wrong in productive ways. It proposes a game that “resembles the real world as much as possible” in order to illuminate the meaning of life, and reaches for Stoicism, Ikigai, socio-economic metabolism, GOAP, and generative agents as the load-bearing materials. The reach is admirable; the load-bearing is not yet load-bearing. What follows is an attempt to walk the same path with shoes on — to keep what is real in the original, name what is wish, and propose a version that could actually be built, played, and survive its own ambition.

§ I.

The four substitution errors

The original document is most useful as a diagnostic. Its weaknesses are the field's weaknesses, repeated with conviction. Four substitutions in particular are worth naming, because every “ambitious sim” pitch — from Spore to Star Citizen to a dozen unshipped Kickstarters — tends to make the same ones.

1. Realism for fidelity

The treatment slips between “realistic” (looks and sounds like the world) and “real” (behaves like the world under stress). These are almost unrelated properties. Dwarf Fortress has the most causally real world ever shipped and is rendered in ASCII. Cyberpunk 2077 at launch looked photographic and behaved like a diorama. A simulator of meaning needs causal depth, not polygon count, and confusing the two is the fastest way to spend $80M on a beautifully-lit theme park.

2. Eudaimonia for an optimization target

Modelling meaning as an Ikigai four-axis tradeoff is the kind of move that sounds rigorous and is, mechanically, a category error. The moment eudaimonia becomes a value the player optimizes, it stops being eudaimonia and becomes another XP bar — subject to Goodhart's Law: when a measure becomes a target, it ceases to be a good measure. The Sims has been doing this for twenty-five years and produced an excellent dollhouse, not a moral teacher. What we actually call a meaningful life is, awkwardly, the residue of choices we did not optimize.

A game that lets you optimize meaning is a game in which meaning is no longer the thing you found.

3. Generative agents for solved technology

The treatment cites LLM-driven NPCs as if Park et al.'s Smallville paper had ended the discussion. It didn't. Smallville showed that 25 GPT-3.5-class agents could plausibly simulate a small town for two days, at a cost the authors politely declined to itemize but that independent estimates have placed in the low thousands of dollars per simulation-day. Scaling that to a world with a thousand inhabitants, running continuously, on a consumer machine, is not a roadmap item. It is the entire research program of a decade. Any serious plan must specify which conversations are LLM-driven (probably: a handful of named characters, on-demand) and which are not (probably: everything else).

4. Project management for game development

The five-phase plan in the original — Planning, Pre-Production, Production, Testing, Launch — describes the surface of game development the way a recipe for “food” describes a kitchen. It is not wrong. It is also not yet useful. What it is missing: budget, team composition, the difference between an AAA path and an auteur path, the failure rate of projects at this scope (it is very high), and any phase-gate criteria for killing the project before it kills the studio. The rest of this essay treats those as the actual hard problems.

§
§ II.

What the field has already proved

The treatment writes as though “a game that resembles the real world” is virgin territory. It isn't. Sixty years of design have already mapped much of it, and any serious project should start by knowing what existing work has demonstrated as feasible — and what it has demonstrated as not.

A working census of relevant prior art and what each title actually establishes
TitleWhat it demonstratedWhat it could not do
Dwarf Fortress (2006–)Multi-century history simulation; deeply coupled physiological, social, and geological systems producing emergent personal narrativesOnboarding; legibility; visual mediation of its own depth
Crusader Kings III (2020)Character psychology, dynastic politics, and stress-driven decay as a generator of moral drama at scaleEconomic depth; ecological coupling; the lives of the unnamed
EVE Online (2003–)A player-driven economy serious enough that CCP hired a real economist (Eyjólfur Guðmundsson) to publish quarterly reports on itOnboarding; meaning outside the spreadsheet
RimWorld (2018)A “storyteller” director-AI architecture that produces narrative arcs from systemic chaos — explicitly designed for emergent stories, not optimal outcomesScale beyond the colony
Outer Wilds (2019)Knowledge-as-progression. The only thing the player gains across the loop is understanding. Mortality is the engineReplayability; scale
Disco Elysium (2019)Skills as competing internal voices; ideology as gameplay; failure as contentProcedurality; the world outside the protagonist's head
Pathologic 2 (2019)Hostile time pressure and irrevocable consequence as moral instruments. The game punishes savingMass appeal
Eco (2018–)Multiplayer environmental governance with a genuine simulated ecosystem and player-legislated lawSolo play; narrative coherence
Kenshi (2018)Faction AI as independent agents pursuing goals without the player; a world that does not revolve around youEmotional resonance
This War of Mine (2014)Care under scarcity as the central mechanic. Moral injury modelled as a stateLong-run civilizational scope
The Sims (2000–)Maslow-style homeostatic needs as a social engine; the everyday as worthy subject matterMortality with weight; meaning that doesn't reset

Two observations follow. First, every component the original treatment proposes — coupled metabolism, agent-based NPCs, eudaimonic mechanics, branching narrative — has been built somewhere. None have been built together. The unsolved problem is not invention, it is composition: how do you put deep ecology, deep economy, deep psychology, and deep narrative in one product without the integration cost devouring all four?

Second, the games that come closest to “exploring the meaning of life” — Outer Wilds, Disco Elysium, Pathologic 2, This War of Mine — are not the simulationally deepest. They are the authorially deepest. This is a fact the treatment should reckon with before it gestures at procedural generation as the route to existential weight.

The deepest games about meaning are not the deepest simulations. They are the most carefully written. This is a fact, not a preference.
§ III.

A realer taxonomy of realism

If “realism” is the load-bearing word, it has to be more specific than the treatment makes it. Four distinct properties get conflated under the same flag, and they impose different costs and yield different payoffs.

Four kinds of realism, and what each actually buys
KindWhat it meansWhat it costs / yields
Sensory realismThe world looks, sounds, and animates like the referentEnormous art budget; yields immersion but not meaning. Diminishing returns past a threshold
Causal realismSystems behave the way the world behaves under perturbation. Cause and effect compose without authored jointsSimulation engineering; yields the “living world” feel. Hardest single property to ship
Phenomenological realismPlaying the game feels, moment to moment, like inhabiting a life. Time pressure, ignorance, embodiment, irreversibilityDesign discipline rather than money; yields existential weight. Often actively hostile to fun
Moral realismChoices carry consequences that do not resolve neatly, do not reward, and cannot be reloaded into oblivionWriting and design courage; yields meaning. Most frequently betrayed by quick-save

The original treatment chases all four and budgets for one (the cheap one, sensory). A defensible plan picks two and goes deep. For a game whose central question is “what does it mean to live a life,” the right two are phenomenological and moral — with just enough causal realism beneath them that the moral weight has a world to land on. Sensory realism is the last priority, not the first.

§
§ IV.

Re-engineering the systems layer

If we keep the treatment's commitment to systemic depth, we have to specify the systems in a way the original does not. Below are the four load-bearing subsystems, each with its honest computational and design budget.

Agent-based modelling at survivable cost

An agent-based simulation is the right substrate for the world. The mature open-source frameworks — Mesa in Python, NetLogo, MASON, Repast — have demonstrated that millions of simple agents can be run on commodity hardware. Real-time, in-game, in a published product, the budget is closer to ten or twenty thousand agents at any one moment, with the rest abstracted into demographic flows. The trick is not running more agents; it is correctly hiding the ones the player can't currently see.

Goal-Oriented Action Planning is appropriate for moment-to-moment behavior — F.E.A.R. shipped it in 2005 — but it is brittle at scale, because the planning cost grows fast with action-set size. Utility AI (used in The Sims and many strategy titles) is cheaper and degrades more gracefully. A realistic stack uses utility AI as the default and reserves GOAP for named characters and combat.

Generative agents, soberly

Park et al.'s Generative Agents paper (Stanford, 2023) is the right reference for what LLM-driven NPCs can do; it is also the right reference for what they cost. The honest budget:

The LLM-NPC budget: what you can and cannot do in 2026
PatternFeasibilityWhy
Every NPC fully LLM-driven, alwaysInfeasibleCost and latency. A town of 200 agents thinking every minute is hundreds of dollars per player-hour even on the cheapest 2026 models
Named NPCs LLM-driven, on demand, with a memory layerFeasibleThis is what Smallville actually demonstrated. Used sparingly, it is shippable
LLM as content-generator at design-time for canned NPC barks and event reactionsAlready standardSeveral 2024–25 titles do this; the LLM never runs at runtime
LLM as referee for ambiguous social outcomesExperimentalThe interesting frontier. Used for “what would this character plausibly do given X”, cached aggressively

Economic simulation as social tissue

The single best existence proof here remains EVE Online, where CCP's decision to publish a Quarterly Economic Report turned an MMO into something economists cited. A realistic plan does not need to replicate EVE; it needs to choose between two well-mapped traditions:

For a game about meaning, the agent-based path is correct — not because it is more accurate (it is and isn't), but because the player must be able to watch a market crash, not be told one happened.

Ecological coupling, honest about abstraction

The original treatment reaches for the nitrogen cycle as illustration. The honest version: a fully process-based nutrient model in the lineage of SWAT (Soil and Water Assessment Tool) or the Madingley whole-Earth ecology model is well beyond a real-time game. But a parametrized version — capturing the topology of nutrient flows, the lag between cause and degradation, and the irreversibility of certain failures — is shippable, and produces the design payoff (long-horizon consequence) without the computational ruin.

The abstraction-fidelity exchange rate

The single most important quantity in simulation design is the ratio of perceived systemic depth to computed systemic depth. Dwarf Fortress's ratio is roughly 1:1 — everything you can see is being computed. The Witcher 3's is closer to 1:0.1 — almost everything is theater. The sweet spot for an ambitious sim is roughly 1:0.3: enough real computation that the world surprises its designers, enough theater that it stays solvent.

The Five-Layer Architecture Authored narrative · storylets · named character arcs LLM-mediated dialogue · on-demand · cached Utility AI default · GOAP for named characters Agent-based economy · bounded rationality · emergent prices Parametrized ecology · lag · irreversibility authored generative procedural emergent substrate heavy writing ms latency, $/call per-frame CPU tick-rate sim background sim
A composition that survives its own ambition: authored layers at top and bottom, generative and procedural in the middle. Cost and risk concentrated where they earn the most weight.
§ V.

Re-engineering the meaning layer

This is the layer the original treatment most needs to revise. Stoicism and Ikigai are real philosophical traditions that have been mistreated by being turned into stat sliders. The argument is not that they can't inform game design; it is that they can't inform it as variables.

What does carry existential weight in games

An audit of the games people actually describe as having changed them — Outer Wilds, Pathologic 2, Disco Elysium, NieR Automata, Spec Ops: The Line, This War of Mine, Shadow of the Colossus, Undertale, Soma — reveals a recurring grammar. None of these games quantify meaning. All of them deploy some combination of:

  1. Irrevocability. Choices that cannot be unmade, even by reload, because the game refuses to honor the save state as moral escape.
  2. Time scarcity. A clock that runs whether the player engages or not (Pathologic, Majora's Mask, Persona's calendar). The opportunity cost of any one action is felt as the action that is foregone.
  3. Opacity of others. NPCs whose interiority is real but not inspectable. The player must infer, not know.
  4. Care under pressure. Systems that put the protagonist's needs against another's, and force the choice (This War of Mine).
  5. The aesthetic of attention. The game rewarding what the player looks at, listens to, returns to — not what they grind (Outer Wilds).
  6. Mortality with weight. Permadeath used sparingly enough that it lands; specifically, not the cheap permadeath of roguelikes, where dying is part of the score loop.
Meaning in games is what the design refuses to quantify. The bar that fills up is the bar that empties the moment.

The Goodhart trap, named directly

The treatment proposes that “balancing the four axes” of Ikigai produces “homeostasis — a stable state of internal well-being.” This is the trap. Once the player can see the four bars, the game becomes about the bars. The activity of finding what you love stops being love-finding and starts being bar-filling. This is not hypothetical — it is what happens, predictably, every time a designer tries it. Apps that gamify mental health show the same failure mode.

The correction is structural. Meaning-related state should be unmeasured by the UI. The game should know what it knows about the player's life and reflect it in the world's response — the way the rumor mill in Crusader Kings reflects the player's choices without ever showing a “reputation” number to optimize against. Authored, opaque, consequential. Not transparent, optimizable, gamified.

Stoicism, returned to its actual claim

The original's Stoic equation — V = f(A, E), where the player's virtuous outcome is some function of internal assent and external event — treats Stoicism as if it were a controller scheme. Its actual claim is harder: that the only stable response to fortune is the cultivation of a disposition that does not require fortune. A game mechanic that captures this isn't a button-press at the moment of crisis. It is the long, accumulated character of the protagonist as expressed in every minor unsurveilled action over fifty hours of play — and revealed, in retrospect, by what the world remembers them for.

This is achievable. It is also expensive: it requires a memory model that is doing real work, an evaluation function that is not legible to the player, and a closing act that has the courage to surface the verdict without offering a redo.

§
§ VI.

Two honest production tracks

There is no single “Ludic Mirror” project. There are two, and they are different products with different audiences and different paths to ruin. Naming them is the first piece of strategic honesty the treatment owes.

Track A (AAA-adjacent) vs. Track B (auteur indie). The middle path does not exist; it is the graveyard of ambitious sims
Track A — AAA-adjacentTrack B — Auteur indie
Budget$35–80M$1–4M
Team size80–150 people peak6–15 people
Timeline5–7 years4–6 years
EngineUnreal 5 or custom; 3D fidelityGodot, Unity, or custom 2D; deliberate visual restraint
Reference titlesRDR2 (scope, not subject); CK3 (depth); Disco Elysium (writing)Dwarf Fortress (scope of simulation); Outer Wilds (scope of meaning); Caves of Qud (scope of strangeness)
Funding modelPublisher or large self-fund. Demands genre legibilityEarly access, patron, or grant. Demands a shippable vertical slice in year 2
Risk profileCatastrophic on miss; pulls a studio underSurvivable on miss; founder takes the personal hit
What kills itScope creep, publisher panic, team attrition at year 4Founder burnout, the “year 5 wall,” community fatigue
Historical precedentSpore (over-promised, under-delivered); Star Citizen (funding without shipping); Death Stranding (rare success)Dwarf Fortress (survived two decades on faith); RimWorld (Sylvester's design discipline); Caves of Qud (small team, fifteen years)

The middle path — $8–20M, 30–50 people, three years — is the path most studios actually attempt and is also, statistically, the one that ships broken or doesn't ship at all. The economics are unforgiving: that budget is large enough that publishers will demand market-tested mechanics, but small enough that the team cannot deliver both the simulation depth and the production polish those demands imply. A serious plan picks a side.

Phase gates with kill criteria

The original's five phases need teeth. A defensible plan attaches a specific survival criterion to each gate — a number or demonstration that, if not achieved, kills the project rather than absorbing more money.

Phase gates that actually gate
GateSurvival criterion
End of vision phase (Mo. 3)A single-paragraph design pillar that ten random people in the studio can paraphrase identically
End of pre-production (Mo. 9)A playable 15-minute vertical slice in which a non-team-member experiences the central existential beat without prompting
Mid-production (Mo. 24)A six-hour playable showing emergent narrative without scripted intervention. If the AI doesn't produce surprise here, it never will
Late production (Mo. 42)A “graveyard test”: external playtesters describe a character's death to their own friends, unprompted, days later
Pre-launch (Mo. 60)Sustained 30+ fps on the median target machine in the worst-case agent simulation. Non-negotiable
§ VII.

The six-month vertical slice

Every ambitious sim project should be capable of producing a six-month prototype that answers, by demonstration, whether the central design conjecture is true. For The Ludic Mirror, the right conjecture is narrow and specific.

The conjecture

That a single protagonist, in a single small town of approximately forty inhabitants, across a single in-game year, can be put through a sequence of small, mostly unwitnessed moral choices whose accumulated weight is felt by the player in the closing minutes as a portrait of who they have become — without any UI ever quantifying that becoming.

What the prototype must contain

What the prototype must not contain

The prototype is not the game. It is the answer to one question: does the central trick work? If at the end of six months a half-dozen external players, asked to describe the experience, return descriptions that converge on the central existential beat — the project is alive. If their descriptions diverge, the conjecture is wrong, and the right move is to kill the project and write a different one. This is what discipline looks like.

§
§ VIII.

Simulation, parable, and the validation problem

The treatment's most ambitious claim is that the game can function as a research instrument — that players can run policy experiments inside it and the results will tell us something true about the world. This deserves a final, careful answer.

They can, and they won't — not in the way the claim implies. A game's simulation is not a model of the world; it is a model of its designers' theory of the world, executed at scale. When players use SimCity to explore urbanism, they are not learning about cities; they are learning about Will Wright's mid-1980s reading of urban economics, refracted through the constraints of a 16-bit machine. This is still valuable. It is not science.

The honest framing is older than the videogame: the simulator is a parable. It tells a story about how its makers think the world is, in a form that the player can interrogate by play. The interrogation is real; the answers it produces are about the parable. To call this a research instrument is to overclaim. To call it useless is to underclaim. It is a tool for moral and conceptual rehearsal, which is what art has always been, and what no spreadsheet has ever been.

The Ludic Mirror does not reflect the world. It reflects a careful argument about the world, and lets the player live inside the argument long enough to notice its shape.

This is the strongest version of the original treatment's vision. The game is not a mirror; it is a lens, ground by particular hands, and what it focuses is a particular question about the meaning of life that its makers thought worth asking. The realer the game gets — in the four senses of real we distinguished in § III — the more weight that question carries when the player lives through it. The aim is not to model the world. The aim is to make an argument the player cannot dismiss.

That is a worthy thing to build. It is also, finally, the kind of thing that takes six years, eighty people or eight, a willingness to ship the unflattering ending, and a refusal of every middle path. The Ludic Mirror, in its first telling, did not yet know that. In its second, it might.

Notes & further reading

  1. Bogost, I. Persuasive Games: The Expressive Power of Videogames. MIT Press, 2007. The canonical treatment of procedural rhetoric; useful corrective to the original treatment's tendency to flatten the concept.
  2. Park, J. S., et al. “Generative Agents: Interactive Simulacra of Human Behavior.” UIST, 2023. The Smallville paper. Read for what it actually demonstrates, not for what excited posters claimed it demonstrated.
  3. Sylvester, T. Designing Games: A Guide to Engineering Experiences. O'Reilly, 2013. The author of RimWorld on the discipline of mechanic-driven narrative.
  4. Adams, T. & Adams, Z. The Dwarf Fortress design talks at GDC across two decades remain the most useful first-hand account of what deep simulation actually costs.
  5. Guðmundsson, E. The CCP Quarterly Economic Reports on EVE Online (2007–2014). The clearest existence proof that a game economy can be a real economy.
  6. Arthur, W. B. Complexity and the Economy. Oxford, 2014. The intellectual foundation for agent-based economic simulation as an alternative to equilibrium models.
  7. The Madingley Model (Harfoot et al., 2014) and SWAT (Soil and Water Assessment Tool) are the right references for what process-based ecology actually requires — mostly so you can decide how much of it you can't afford.
  8. Hocking, C. “Ludonarrative dissonance in Bioshock.” Click Nothing, 2007. The original diagnosis of the gap between what a game says and what it makes you do; the underlying problem any “meaning of life” sim must solve.
Composed in response · Eight movements · A re-engineering
Set in Cormorant Garamond & Source Serif · Printed in cream & oxblood