Running an LLM-driven town with 800+ persistent agents: concurrency, context caching, and inference costs

Wait 5 sec.

I spent the past year independently building Slow Vale, an LLM-driven life simulation. The Chinese server now has 800+ AI residents sharing one continuously running city. This is an engineering write-up about concurrent decisions, dynamic action spaces, context caching, and the operating costs of a persistent multi-agent system. The runtime currently uses hosted DeepSeek Flash, rather than local inference. I am the developer. I wrote the original material in Chinese and used AI to translate and refine the English. Product metrics below are current through October 7, 2026. Asynchronous decisions in a continuously advancing world Each character makes roughly 300–400 LLM calls per day, with an average context of around 30,000 tokens per call. A call includes the character's state, relevant experiences, current environment, and available actions. The model selects an action and its parameters; the backend turns that decision into an activity that occupies time and resources. Game time and real time coexist. Sleeping might occupy 8 in-game hours, while saying one sentence might take 1 in-game minute. Inference itself takes real time. While a call is in flight, other characters can change the environment, and the world clock continues to advance. Interactions also involve mutual exclusion. If A is talking with B, C cannot simultaneously pull B into a separate conversation. Facilities, production tasks, and other activities have their own rules for acquiring and releasing occupied resources. Separating concurrent inference from world-state mutation LLM calls can run concurrently, but model responses do not directly mutate the world. Results return to the world's execution flow, undergo validity checks, and are applied by the execution component that owns world state. For example, the last fish on a shelf might still be available when a character starts inference. By the time the response arrives, another resident may have bought it. The purchase intent must be checked against current inventory. Similarly, the person a model wants to talk to may have left, gone to sleep, or started another activity. There is therefore an explicit time gap between the context used for a decision and the state at execution. The system must distinguish what a character intends to do, whether the action is still valid, and what effects have actually occurred. Completion, failure, interruption, and recovery each need consistent state transitions. A shared runtime for activities that occupy time Movement, production, conversation, and sleep have different durations, participants, and completion conditions. A common runtime makes it possible to manage busy characters, resource conflicts, and service recovery without building a separate scheduler for every mechanic. The frontend must also follow actual progress: when an activity started, how long it has been running, whether it completed, and what it produced. Logs and scene animations need to correspond to facts committed by the backend. This is a significant source of complexity in a persistent world: one event can affect future decisions, persistence, other residents, and the player interface. Decision context is part of the backend architecture A personality description alone is insufficient for a character that acts over long periods. Each decision needs the character's current needs, location, assets, ongoing concerns, relevant relationships, and the actions actually available at that moment. These inputs have different update frequencies and lifetimes. Personality is relatively stable; hunger and energy change continuously; inventory and other characters' states can change within seconds. An experience may continue to affect a relationship long afterward. Each kind of information needs rules for entering context, updating, and leaving the character's current attention. The action space also needs to reflect game state. Options presented to the model should disclose their execution conditions and relevant state, while the backend retains final validation. Otherwise, characters repeatedly attempt unavailable actions or spend calls trying to understand rules that were never clearly disclosed. I have invested substantial effort here: organizing stable and dynamic information, controlling irrelevant history growth, avoiding duplicate reminders, and keeping context prefixes stable. This affects behavior quality, inference latency, and cache hit rates, making it part of the backend architecture. Roughly 5 billion tokens a day for under $100 in model fees The Chinese server currently processes around 5 billion tokens per day, with model fees below US$100. It primarily uses inexpensive models such as DeepSeek Flash, while maintaining a cache hit rate above 90%. The token count includes cached input. In a system with frequent calls, many characters, and long contexts, reusable stable prefixes directly affect the bill. Which information stays stable, which changes on each call, and how it is ordered all require deliberate design. https://preview.redd.it/71a7tjghs7uh1.jpg?width=1360&format=pjpg&auto=webp&s=5a57fc15627240d8a416b48947d9b6bd34eed474 Actual DeepSeek usage and billing for October 7, 2026 (GMT+8): approximately 4.424 billion tokens and 163,742 requests across all API keys, costing CNY 472.33. The model shown for that day is deepseek-flash. The trade-off between dynamic action spaces and prefix caching One concrete engineering trade-off was how to represent a dynamic action space when using tool calling or structured output. The available actions and parameter values change on every decision: which facilities are nearby, which goods are available, and whom the character can talk to all depend on the current world state. Encoding these options directly in tool definitions or an output schema gives stronger output constraints, but also makes the schema change frequently. In some API implementations I tested early on, those definitions became part of the request prefix. Changing the schema prevented the otherwise stable context after it from hitting the cache. At that stage, I chose ordinary text generation of JSON for the primary path, with parsing and validation in the backend and a strict-schema fallback when parsing failed. The model still received explicit, state-dependent action options, but those options lived in the current decision context rather than in a changing output schema. This kept stable instructions and reusable history toward the front, with current state and action options toward the end. The trade-off was giving up decoding-time format guarantees on the primary path. The application had to handle malformed output and validate actions and parameters against the world state at execution time. The 90%+ cache hit rate therefore comes from designing the whole request structure, rather than simply enabling a provider feature. The percentage refers to the share of input tokens served from cache; the model still generates a fresh output for every decision. When comparing invocation modes, I consider format reliability, character behavior, cache reuse, latency, and cost together. These figures cover model fees. As the resident population grows, database load, state delivery, log storage, and scene rendering also matter. Inexpensive inference makes continuous simulation feasible; sustained operation still depends on resource management across the entire system. Organizing AI collaboration with runbooks Maintaining this many modules alone requires giving AI a reasonably complete working environment. I provide development and operational tools, including access to logs, Langfuse, growth analytics, the database, and procedures for maintaining production services. https://preview.redd.it/iqo7313ks7uh1.png?width=962&format=png&auto=webp&s=a6af26e071e8b76ff55c0d708a223e53664a89e8 My Codex usage: approximately 43.25 billion cumulative tokens and an 85-day longest streak. Codex is only part of the AI coding tooling I use. These are development usage figures, separate from the model calls that power the game's residents. A set of runbooks governs their use. The project has extensive documentation, organized by task and module. It specifies which documents must be read for each task, which sources define current contracts, which decisions only I can make, and which documents AI should maintain when it discovers drift from the implementation. Task entry points and action boundaries are central. An investigation starts by identifying the data source and time window. Permission to query does not imply permission to modify production data, and permission to fix code does not imply permission to deploy it. Access to a tool needs to come with explicit conditions for using it. I have also turned recurring maintenance into automated workflows: diagnosing and fixing production problems, daily in-depth reviews of character behavior and gameplay outcomes, and daily cleanup of maintenance code that has served its purpose. Each workflow specifies the evidence required, permitted actions, validation, and stopping conditions. My involvement varies by area. I directly decide or closely participate in frontend/backend contracts, backend architecture, and ownership of state and resources. For frontend and Phaser implementation, I focus more on evaluating the result, while still defining design tokens, page structure, reusable components, and presentation boundaries. This approach depends on maintainable project knowledge. Constraints discovered during a task need to return to the formal documentation, and outdated procedures need correction. Otherwise, as the project grows, AI can implement a locally plausible change based on old assumptions while breaking contracts elsewhere. Three to five production releases a day The city has been running for more than 350 in-game days, equivalent to nearly 100 real-world days. A substantial portion of the earliest players are still playing. I built the entire project myself, including the backend, frontend, Phaser scenes, content production, monitoring, and operations. It now contains more than 400,000 lines of code, including over 200,000 in the core backend, across approximately 2,200 commits. I use AI coding tools extensively. I make or closely participate in decisions about product direction, core mechanics, and architectural boundaries, while AI handles much of the implementation, investigation, and maintenance. As the project moved from a prototype to a continuously operating product, system design and the development workflow became a major part of the work. I currently deploy an average of three to five times a day. Releases include architectural changes, balance and gameplay adjustments, new systems, UI and art changes, performance improvements, and bug fixes. The project has approximately 2,200 commits, with more than ten commits per day during active development. The iteration speed comes from a short feedback cycle between implementation, observation, and adjustment. Players continue to inhabit the same city. After a feature goes live, I can observe actual usage and character behavior, then decide whether to change a mechanic, clarify what information characters receive, or fix an implementation issue. I assess software operation and gameplay outcomes separately. Error rates, latency, database load, and model calls indicate whether the system is operating normally. Understanding whether characters repeat themselves, understand a new mechanic, or successfully complete production and social activities requires reading their actual experiences and decision traces. Monitoring, queries, behavior evaluation, and repair workflows are therefore part of daily development. Frequent releases also require clear module boundaries, validation scope, and recovery procedures, along with prompt removal of temporary maintenance code. Otherwise, fast individual changes can still make the system progressively harder to maintain. From a virtual pet to hours of viewing I initially imagined the game as a kind of virtual pet. Players would open it once a day, check that their character had eaten and earned some money, perhaps send a message, and leave. A different pattern emerged in actual use. Some players watch it like a livestream, spending several hours a day observing their character. They follow the progress of a relationship, check whether a shop has customers, or wait to see whether the character follows a suggestion they just sent. The product therefore needs to support both brief check-ins and continuous viewing. Over the past month, daily active users on the Chinese server grew from 177 on September 10 to 865 on October 7, approximately 4.9 times the starting figure. Between October 1 and October 7, DAU grew from 431 to 865. Growth during this period came primarily through players sharing the game organically. https://preview.redd.it/c4ne6icks7uh1.png?width=2000&format=png&auto=webp&s=13ea0fc510e42563dbce937bd8485506beca56cf Chinese-server DAU, measured as distinct users who successfully entered the game. Chart redrawn from PostHog query results; dates use Asia/Shanghai. For the 225 users who first successfully entered the game in August, exact-day retention was 68.9% on Day 1, 56.0% on Day 7, and 40.9% on Day 30 (155, 126, and 92 returning users). On October 7, the 850 non-admin users with valid foreground-duration records had a median of 29.9 minutes and a P90 of approximately 4 hours. During October 1–7, 86 users were active on at least four days and averaged at least three foreground hours per active day. https://preview.redd.it/ixfcg1nks7uh1.png?width=2000&format=png&auto=webp&s=c7be00205bfdc85a636b0dcdeb246f081dae7485 Retention for the same cohort of 225 first-time entrants: 155, 126, and 92 returning users, respectively. Foreground usage measures time with the game in the foreground; it does not establish uninterrupted attention. Together with player feedback, it indicates a stable group of users who spend long periods with the game. This creates specific engineering requirements. Occasional visitors need to understand what happened while they were away. Continuous viewers need to see activities progress, understand why a character acts, what they are waiting for, and how an interaction ends. Activity logs, recaps, and live scenes are all core interfaces. More than 800 residents sharing one city https://preview.redd.it/rxxyxgvls7uh1.jpg?width=1080&format=pjpg&auto=webp&s=cd20e91611d97716981186de58c9cfa84561830a The city. Shops and workplaces in the shared environment support actual game activities. Players create a character with a personality of their own, influence them through messages and gifts, and observe their life. LLMs decide the character's movements, meals, sleep, work, and social interactions. Characters created by other players inhabit the same world. They can talk in real time, trade, share meals, fall in love, and live together. Residents need to earn a living. They can run farms and ranches, fish by the sea, work in an office, open their own shops, or sell goods at a market stall. These activities connect to a shared economy: residents produce agricultural goods, products have actual inventory, supply and demand affect prices, and business owners bear costs and make purchasing and pricing decisions. All food is produced through residents' labor. Restaurant owners manage their businesses, cooks prepare meals, couriers deliver orders, and customers pay for and consume the food. Each meal has a chain of ingredients, production, service, and consumption behind it, with city residents participating at every stage. https://preview.redd.it/82xpv54ms7uh1.png?width=1079&format=png&auto=webp&s=ec97f5fc5a6f9890fa86a3bd9ab11e00d431f361 Farming and ranching. These four English showcase images use the game's native renderer and UI with staged scenes and demonstration data. https://preview.redd.it/4kyosd64t7uh1.png?width=1079&format=png&auto=webp&s=d6126c8e0e6b368879b8c213db375c16272fc99c A resident sowing seeds. Phaser scenes visualize everyday production activities. https://preview.redd.it/cj791fe6t7uh1.png?width=860&format=png&auto=webp&s=73075b273bfd1848dc00842db704374c359c569d Farm management: crop growth, livestock, feed, and production status. https://preview.redd.it/8siqix08t7uh1.png?width=860&format=png&auto=webp&s=5c80c6810b42123aa97c9bac5c2f6cfe532fca90 Market inventory, resident shops, and price trends. The values shown here are demonstration data. Players can view live scenes, character status, relationships, and activity logs, and receive postcards from their characters. Relationships accumulate through interactions that actually take place. Events from a character's life become part of the context for later decisions. Dreaming is a recent addition. While sleeping, characters generate dreams based on their experiences, and occasionally talk in their sleep. For example, Bread Pitt on the English server dreamed that a courier was chasing him down an office hallway with a burger he had already paid for. Every door led back to two friends who were somehow still hungry. In his sleep, he muttered: “Just leave it at the door…” https://preview.redd.it/knp8qui9t7uh1.jpg?width=1220&format=pjpg&auto=webp&s=36fade0de8c93ce0b4e72067a4f183d28d39646c An actual dream from the English server. Dreams and sleep talking appear in the sleep activity log, using the existing decision and logging mechanisms. These details give players a sense of continuity in the character's life. A day's work, friends, or a missed meal can reappear in a different form in later experiences. Engineering for a persistent world The project has grown from a character prototype into a continuously running city. Residents share inventory, facilities, space, and time. Their actions change the conditions for other characters' next decisions. An action produced by inference must remain valid in the current world and survive persistence, delivery to the interface, and service recovery. Player behavior is also changing my understanding of the product. It can be a virtual pet checked once a day, or a life simulation watched for hours. Long-term players accumulate knowledge of characters, relationships, and the city, making continuity an important part of the experience itself. I will continue improving the mechanics, presentation, and scalability of this persistent world. The English browser version is available at slowvale.com. No invitation code is required; you can register with an email address and start playing.   submitted by   /u/Low_Bad_6585 [link]   [comments]