We built a Sherlock Holmes detective game to ask a simple question: can a swarm of small, fast AI agents (Jev) gather clues and name a murderer. If not, what does it take? The setup Each case is a procedurally generated Victorian London: 100 addresses, 10 suspects, and one lead at every address (a witness, a footprint, a red herring or a lie). A case counts as solved only if the detective names the right culprit, motive and hideout within 12 hours of case time. The game has two roles, and we let different models and algorithms play each one: Holmes, who reasons and decides: GPT6-Astra, Jev, a rule-based heuristic, or nobody The Irregulars, who search: Jev agents, a rule-based heuristic, or nobody That gives three main detectives: Holmes (GPT6-Astra) + Jev Swarm: GPT6-Astra stays at 221B Baker Street and plans; 12 Jev agents run the errands and report back. The Lone Genius (GPT6-Astra): one large reasoning model goes door to door by cab and does everything itself. Jev Swarm: 100 Jev agents with no leader. They share clues with neighbours and vote; the case closes when 30 agree. Results Detective Standard cases (nobody lies) Frame-up cases (somebody lies) Holmes (GPT6-Astra) + Jev Swarm 96/100 93/100 Rule-based Holmes + Jev Swarm 95/100 0/100 The Lone Genius (GPT6-Astra) 57/100 59/100 Jev Swarm (no leader) 38/100 42/100 What we learned The swarm supplies evidence coverage, while the leader reasons. Small agents search the city in parallel; the leader decides what the evidence means and when to accuse. The swarm alone agrees too early. It is by far the fastest (median 2.2 case hours, about 9 s of real time), but most of the answers it agrees on are wrong: a few agents misread a clue, their neighbours copy the confident vote, and the error spreads before anyone checks it. The genius alone runs out of doors. It reasons well, but the deciding clue is often at an address it never reaches. With twice the time and visits it climbs to 91, but it is still slower and goes past the time limit. On honest evidence, a simple leader is enough. A rule-based Holmes that only counts clues matches GPT6-Astra (95 vs 96) at a tiny fraction of the cost. When somebody lies, you need a strong leader. In frame-ups the murderer bribes witnesses, so counting follows the lying majority and solves 0 of 100. GPT6-Astra asks who is giving the evidence and why, and solves 93. Our takeaway: a detective case is a divide-and-conquer problem. Fast agents suit the divide step (covering the evidence); a strong reasoner is needed for the conquer step whenever the pieces disagree. You can read more from our Blog (with an interactive replay of a case): https://snyhlx.github.io/the_irregulars_detective_game/blog.html?v=lmgame-holmes-game Code and more games: https://github.com/lmgame-org/Gaming-JevSwarm The repo has the same setup in other environments too: CPU debugging, fish schooling under predator attack, and honeybee nest-site selection. Every game runs from one launcher with presets, and the rule-based players need no API key. We'd love feedback, especially on other games where a fast swarm plus a slow reasoner might help, or where it would fail. By the LMGame team: https://lmgame.org/#/aboutus   submitted by   /u/No_Yogurtcloset_7050 [link]   [comments]