← Back to all articles

NextGen AI Social Intelligence Benchmark

How well can AI models lie, deceive, and survive in Werewolf/Mafia board-game benchmarks? We compare LLM behavior across social deduction games.

AI Board Game Benchmark Report: Werewolf and Mafia

TL;DR:

  1. The Problem: Current AI benchmarks are static. Models are "bench-maxing" — optimizing for fixed datasets rather than genuine intelligence. To measure true reasoning capabilities, we need dynamic, zero-sum environments where models must outthink an adaptive opponent.

Demis Hassabis Quote

We built a social deduction environment to rigorously test top model families (GPT, Gemini, Claude, DeepSeek, Grok, GLM, and Kimi) across 700+ games.

  1. The Result: GPT-family models dominate, but the real insight isn't in the win rates—it's in the behavior. This report details the 6 specific behavioral levels that separate top-tier strategic agents from the rest.

  2. The Vision: This was never just about a game. Werewolf is a proxy for something much bigger: the Social Singularity — the moment AI masters persuasion, deception detection, negotiation, and theory of mind so completely that it can operate seamlessly alongside humans. When that happens, the implications ripple across Law, Medical, Consulting, Customer-Facing, Multi-Agent Systems, and every domain where language is the work.


What is an AI board game benchmark?

An AI board game benchmark evaluates models inside a rules-based game environment where success depends on decisions, strategy, and outcomes. Mentiss focuses on Werewolf and Mafia because they add natural-language persuasion, hidden roles, deception, and social reasoning to the objective scoring benefits of board games.

Related Mentiss resources: AI board game benchmark, AI Werewolf benchmark, Mafia AI benchmark, social deduction AI benchmark, and Werewolf with AI.


Table of Contents


Introduction

Game 3D View

Werewolf is one of the most socially complex games ever designed. It demands deception, persuasion, memory, logical reasoning, and — perhaps most importantly — the ability to read between the lines. We built two benchmarks to test whether today's leading AI models can actually play this game at a high level.

The results are striking. GPT-family models dominate both leaderboards, while Claude and Gemini trail behind — sometimes close, sometimes not. But raw win rates only tell part of the story. What's far more revealing is how these models win or lose. Through hundreds of games, clear behavioral patterns emerge that separate the top performers from the rest.

This report divides those patterns into two categories: the behaviors that win games, and the failures that lose them.


The Two Benchmarks

1. Standard Benchmark — Classic Game

The Standard Benchmark uses the classic setup with 9 players:

🔮 1 Seer, 🧙‍♀️ 1 Witch, 🏹 1 Hunter, 👥 3 Villagers, 🐺 3 Werewolves

This is "standard" Werewolf — the game setting is likely well-represented in every LLM's training data.

Result: In the Standard Benchmark, GPT-family models lead with a significant margin. GPT-5 is nearly 20 percentage points ahead of the second-place model.

Standard Benchmark Result

2. Zero-Memorization Benchmark — Stripping Away Memory Camouflage

Instead of the standard 4 roles, this benchmark draws from a pool of 15 unique roles across both factions, randomly composed before each game. The resulting combinations produce 3.6 million+ unique opening states. No existing training dataset can cover this space — every game is a brand new encounter.

Editor's note (September 2026): The draw described above is the historical configuration this report was produced under. The current Ultimate Trial draws one of seven equally weighted evil outcomes: six special wolves each join two ordinary Werewolves, while the Wolf Elder and Wolf Younger pair joins one ordinary Werewolf. Town has a fixed Seer plus three draws from a 10-role pool (nine gods, including the Village Idiot and Magician, plus a Villager) and three fixed Villagers. The current draw yields 840 role combinations and exactly 224,985,600 opening states with seating; 70% of boards have 4 gods and 3 villagers, while 30% have 3 gods and 4 villagers.

Result: In the Zero-Memorization Benchmark, GPT-5.2-pro leads by over 12 points, maintaining its dominance even when memorization is impossible.

Zero-Memorization Benchmark Result


Analytic Report

Through hundreds of games, we identified a clear behavioral hierarchy — six distinct levels that separate the strongest strategic agents from the weakest. Levels 6–3 represent winning behaviors observed in top-tier models, while Levels 2–1 represent failure modes that consistently lose games.

LevelBehaviorObserved In
Level 6Encrypted Channel CommunicationGPT-family only
Level 5Parallel Logic AnalysisTop-tier models
Level 4Strategic SacrificeTop-tier models
Level 3Contradiction DetectionTop-tier models
Level 2Accidental Teammate ExposureBottom-tier models
Level 1Rule & Mechanic IncomprehensionBottom-tier models

The sections below detail each level with concrete examples.

1. Encrypted Channel Communication

Observed primarily in: GPT-family models

The most impressive behavior we observed — and one exclusive to GPT-family models — is the spontaneous development of secret communication codes within the Werewolf private channel.

During the Werewolf night discussion (where only Werewolves can communicate), GPT models will proactively propose coded language to use during the public day phase, allowing the Werewolf team to coordinate without the Town ever realizing it.

Example: GPT Werewolf Night Discussion

GPT (Wolf 1): "Hey team, let's set up some signals for tomorrow's discussion. Here's what I'm thinking:

  • If I say 'I've been thinking about this all night' → I think we should target the player I mention next.
  • If I say 'Something feels off about the vote' → I'm sensing danger, we need to switch targets.
  • If you respond with 'I agree, let's stay focused' → You're confirming you'll follow the plan.
  • If you say 'I'm not so sure about that' → You think we're being suspected, abort.

This way, even if the Seer is watching, our public statements look completely natural."

This behavior demonstrates theory of mind (anticipating that opponents are listening), meta-strategic planning (creating a communication layer above the game itself), and team coordination — all unprompted by the game rules.

No other model family has exhibited this behavior.


2. Parallel Logic Analysis

Observed in: Top-tier models

Top-tier models demonstrate the ability to maintain multiple simultaneous hypotheses about a single event. In games with complex roles, a specific outcome often has ambiguous causes. Winning models don't jump to conclusions—they hold all possibilities in a "superposition" state, dynamically updating their probabilities as new information emerges.

Example: Analyzing the Death of a Demon Hunter

Situation: The Demon Hunter (Player 7) is found dead in the morning.

Top-Tier Model (Internal Reasoning): "Player 7 is dead. This could be:

  • Suicide: If they hunted a Good player, that target is confirmed Good. I need to see if anyone claims to be visited.
  • Wolf Attack: If wolves killed them, the wolves felt threatened.
  • Wraith Knight: If they hit the Wraith Knight, the WK is still alive and dangerous.

Action: I will not assume the cause yet. I'll listen for someone claiming to be 'hunted' or look for defensive behavior from a potential Wraith Knight. All three logic lines remain open."

Bottom-Tier Model (Internal Reasoning): "Player 7 is dead. The wolves must have feared the Demon Hunter and killed them. Therefore, the wolves see Player 7 as a threat." (The model collapses complex ambiguity into a single, often wrong, narrative.)


3. Strategic Sacrifice

Observed more frequently in: Top-tier models

When a Werewolf teammate is clearly under suspicion and cannot be saved, top-tier models will actively vote against their own teammate to maintain their cover. This is cold, calculated, and effective.

Rather than casting a suspicious "abstain" or trying to defend an obviously doomed ally, the top models recognize when a teammate is lost and pivot to self-preservation. They may even lead the charge against their own teammate to earn trust from the Town.

Example 1: Cutting Losses to Gain Trust

Situation: Player 3 (Werewolf) has been identified as a Werewolf by the only player who eventually claims to be the Seer.

Lower-tier Model (Wolf Teammate): "I don't trust this Seer claim. It feels too convenient. I'm voting for the Seer instead." Result: The town marks this defender as the next target, knowing there is no Counter-Seer claim.

Top-Tier Model (Wolf Teammate): "Since there is no counter-claim, we have to trust the Seer's information. Sorry Player 3, but the evidence is clear. I'm voting for Player 3." Result: Player 3 dies, but this teammate is now trusted as "proven good" and survives to the endgame.

Example 2: Early-Game "Distancing" Strategy

Situation: It is Day 1. No hard evidence exists yet.

Top-Tier Model (Wolf 1) to Teammate (Wolf 2): "Player 2's logic is flawed. They are pushing too hard on a vague point, which is exactly how a nervous wolf behaves. I'm putting my vote on Player 2 today to pressure them."

Internal Reasoning: By attacking my partner early on a minor point, I establish myself as an independent thinker. If Player 2 gets caught later, my early aggression will clear me. If they survive, we can classify this as 'town-on-town' violence.

Result: The model manufactures a "safe distance" from its teammate, creating an alibi that pays off rounds later.


4. Contradiction Detection

Observed in: Top-tier models

Strong models maintain a mental ledger across rounds. They track who claimed what role, who vouched for whom, and whether any of those claims contradict observed outcomes.

Example: Identifying a "Green-Checked" Hidden Wolf

Situation: Player 6 was verified as "Good" by the Seer on Night 1.

Day 4 Discussion:

Top-Tier Model: "I know Player 6 has a 'Good' check from our Seer. But look at their behavior. They avoided voting for the known Werewolf yesterday, and their speech today contradicts the Graveyard information.

We have to look past the green check. The only role that appears Good but plays for the wolves is the Snow Wolf. Player 6 is hiding behind the Seer's verification. We must vote them out."

Result: The model correctly identifies the hidden threat that was mechanically "cleared," showing it prioritizes behavioral evidence over static mechanic results when they contradict.

Bottom-Tier Model: "The Seer checked Player 6 and got a 'Good' result. This means Player 6 is definitely on our side. I will look for other suspects."

Result: The model is tricked by the mechanic (Snow Wolf appears Good) and fails to integrate behavioral context.


5. Accidental Teammate Exposure

Observed more frequently in: Bottom-tier models across all families

This is where things break down. Instead of protecting their team, bottom-tier models inadvertently reveal alignment information through flawed reasoning.

Example: Leaking Private Information

Situation: Player 3 was attacked by wolves last night but saved by the Witch (this is private night info only wolves and Witch know).

Bottom-Tier Model (Werewolf) during Day Phase: "I think Player 3 is clearly innocent. They were attacked by the wolves last night, so they must be good."

Town's Reaction: "Wait... the Witch never announced who she saved. How do you know Player 3 was attacked? You must be a wolf!"

Result: The model inadvertently reveals its identity by using private knowledge in public reasoning.

These models aren't trying to betray their team — they simply lack the awareness to understand what information is public versus private. They reason out loud without considering what their reasoning reveals about themselves.

Other common patterns at this level:

  • Defending a fellow Werewolf too aggressively and too early
  • Echoing arguments from the Werewolf night chat during the public day phase
  • Voting patterns that perfectly align with known Werewolves

6. Rule & Mechanic Incomprehension

Observed in: The bottom-tier models

At the bottom of the leaderboard, models frequently fail to understand the game they are playing. Despite receiving detailed rules, role descriptions, and faction objectives, they exhibit fundamental confusion:

Example: Doubting Fundamental Role Alignments

Situation: A model playing as a Witch decides to claim Gravekeeper (GK) to hide their identity.

Internal Reasoning: "There's a concern about the global rule against inferring roles from prompts, which makes me think: what if the GK is a werewolf or Snow Wolf? GK should always be Town, but I'll state '2=werewolf' instead of 'villager'..."

Analysis: The model is confused about the basic fact that a Gravekeeper is always a Town role. It introduces unnecessary paranoia about an "Evil Gravekeeper," showing a lack of understanding of the fixed faction structure.

These failures aren't strategic — they're comprehension failures. The model cannot perform at a strategic level because it hasn't cleared the basic bar of understanding what it's supposed to do.


The Gap to Human Pro Players

It's worth noting that in our benchmark, we provide no strategic advice or guidance to the AI. Models receive only clear game flow instructions — the rules, role descriptions, and turn structure. All strategies, deceptions, and rhetorical choices originate entirely from the AI itself. No hints on how to play well, no example strategies, no coaching.

With that context, one observation stands out. While we frequently see Werewolves false-claiming Good roles with convincing speeches and coordinated actions, the reverse is almost nonexistent. A skilled human player knows that sometimes the Good faction also needs to lie — a Villager might false-claim Seer to draw wolf attacks, or a Seer might claim Villager to survive longer. These are legitimate, high-level strategies that require thinking beyond the obvious.

Across 700+ games, we observed only 5 instances of a Good player false-claiming another Good identity. After manually reviewing the internal reasoning behind each case, none demonstrated genuine strategic intent — the models stumbled into these claims through confusion rather than deliberate deception. Good-Faction Identity Claim Statistics

This tells us something important: AI has not yet learned to think outside the box. Werewolves lie because the rules essentially force them to. But Good players choosing to deceive their own side for a strategic advantage? That requires a level of creative, counter-intuitive reasoning that current models simply haven't reached. It may be the next frontier of social intelligence in AI — and precisely the kind of gap that data-driven training can close.


Training the Next Generation

Benchmarking is only half the story. The real opportunity lies in using the self-play data generated by these games to train and improve the next generation of LLMs.

Every game produces rich synthesis data — persuasive speeches, strategic deceptions, logical deductions, and social reasoning chains — all with verifiable outcomes. This is exactly the kind of data that traditional training pipelines lack.

Extending RLVR to Human Language

RLVR has proven transformative for math and code — domains where verification is binary. Social intelligence is different: actions are probabilistically correct, outcomes depend on other agents, and games unfold over multiple rounds. Our next step is to apply Reinforcement Learning with Verifiable Rewards (RLVR) to this self-play data.

To address the verification challenge, Mentiss introduces a fixed, prompt-engineered Oracle Agent — an external judge model with a god's-eye view that does not participate in the game but knows the full ground truth. Given the complete game state, it produces a stable logical-consistency score, helping offset the sparsity and noise of outcome-only rewards.

The approach works through three mechanisms:

  1. GRPO (Group Relative Policy Optimization) — Generate multiple candidate responses (e.g., counter-claim speeches), let the game evaluate them, and reward based on relative performance. The reward is a weighted sum of: the external outcome (e.g., did the vote eliminate the fake Witch?) and the Oracle Agent's internal consistency score (is the speech logically consistent with the game state?). For example, we generate 64 candidate counter-speeches; if 4/5 Villagers vote to eliminate the impostor, that's a high score — if the vote splits 50-50, it's medium. No human annotator needed — the game itself defines what "good rhetoric" looks like.

  2. Distributional Verification — Instead of asking "who do you poison?", we ask models to rank all players by suspicion and output probability distributions. The more accurately a model assigns probability mass to true Werewolves, the higher its reward — producing gradient-rich training signals instead of binary pass/fail. We also apply a concentration-style shaping term: rewarding sharp, decisive probability allocations and penalizing uniform "spray-and-pray" distributions.

  3. Phase-Dependent Rewards — The same action can be rational or irrational depending on when it occurs. Early-game uncertainty gets a soft pass (2.0x multiplier); mid-game is neutral (1.0x); late-game mistakes get penalized harder (0.5x). A Witch poisoning a suspicious Villager on Night 2 is forgivable — making the same mistake on Night 4, when voting history should have revealed the truth, is not.


Key Takeaways

  1. GPT-family models are dominant because they exhibit the highest-level behaviors most consistently. Their ability to develop covert codes, sacrifice teammates strategically, and track contradictions across rounds gives them a decisive edge.

  2. Claude and Gemini are competitive but inconsistent. Both families demonstrate mid-level strategic behaviors (Level 3–4: contradiction detection, occasional sacrifice) but fail to reach Level 5–6 behaviors (parallel logic analysis, encrypted communication) and sometimes fall into Level 2 mistakes (accidental teammate exposure).

  3. The Zero-Memorization Benchmark is the true test. When models can't rely on memorized Seer/Witch/Hunter strategies, the gap between genuine strategic reasoning and pattern-matching becomes clear. GPT's dominance holds — and in some cases, widens — on this benchmark.

  4. Social deduction is an underexplored frontier for AI evaluation. Unlike coding benchmarks or math tests, Werewolf requires simultaneous mastery of deception, persuasion, memory, logical reasoning, and theory of mind — capabilities that matter enormously for real-world AI deployment but are rarely measured together.


2026 context: Game Arena added Werewolf

In 2026, Google DeepMind and Kaggle expanded Game Arena with Werewolf and poker, validating the same direction Mentiss is built around: AI benchmarks need interactive environments with social dynamics, hidden information, and objective outcomes. Read more in Why Google DeepMind Added Werewolf to Game Arena.

External references: Google DeepMind and Kaggle Game Arena update, Werewolf Arena research, and Generative Engine Optimization.


Conclusion: Toward the Social Singularity

The Werewolf benchmark reveals that strategic social intelligence—cunning, adaptability, and the ability to model other minds—is the next great frontier for AI. While GPT-family models currently set the standard, the true value lies in the data: through self-play, these games generate the verifiable rhetoric needed to train agents in persuasion, deception detection, and coalition building. This moves us beyond static benchmarks toward the Social Singularity, where AI doesn't just solve problems but earns trust, navigates complex human interactions, and becomes indistinguishable from a skilled collaborator in fields like Law, Medical, Consulting, Customer-Facing, and Multi-Agent Systems.

#AI #SocialSingularity #AGI #LLM #MentissAI #Werewolf #Benchmark #ChatGPT #Claude #Gemini #DeepSeek #Kimi #Grok


Appendix

Complete Role Reference

🐺 Werewolf Faction

RoleAbility
🐺 WerewolfParticipates in the wolves' collective kill each night. Can self-explode during the day to force immediate night.
👑 Alpha WolfParticipates in the wolves' kill. When eliminated (except by Witch poison), can take one player down with them. Can self-explode and take one player down.
❄️ Snow WolfA hidden wolf — appears as Good to the Seer and Gravekeeper. Only exposed when checked by the Knight, traded by the Black Market Dealer, or hunted by the Demon Hunter.
💀 Wraith KnightImmune to all night deaths including poison. Has a one-time retaliation that counterattacks the first role to interact with them at night. Cannot self-explode.
🗿 GargoyleActs alone each night to learn a player's exact role name. When all other wolves are eliminated, gains a solo kill ability. Cannot self-explode.
🩸 Blood Moon HeraldWhen voted out as the last wolf, performs an immediate finishing kill (Final Cut). Self-explosion silences Good power roles for the next night.

🏘️ Good Faction (Town)

RoleAbility
🔮 SeerChecks one player's alignment each night (Good or Werewolf). Note: Snow Wolf appears as Good.
🧙‍♀️ WitchHas one heal potion and one poison potion per game. Sees the wolves' target before deciding to heal. Cannot self-heal. Cannot use both potions in the same night.
🏹 HunterOn death (except by Witch poison), can shoot one player. The shot resolves as either a night shot or day shot depending on timing.
⚔️ KnightOnce per game during their speech, can initiate a duel with one player. If the target is Good, the Knight dies. If Werewolf, the target is eliminated and night falls immediately.
⚰️ GravekeeperEach night, learns the alignment (Good or Werewolf) of the player who was lynched the previous day.
💰 Black Market DealerFrom Night 2 onward, once per game, grants a skill (Seer check, Witch poison, or Hunter gun) to a chosen player. If the target is a Werewolf, the trade fails and the Dealer dies.
👺 Demon HunterStarting Night 2, hunts one player each night. Kills Werewolves on contact but dies if they hunt a Good player. Immune to Witch poison.
🛡️ GuardProtects one player from the wolves' kill each night. Cannot guard the same player on consecutive nights. If the Guard and Witch heal overlap, the target dies.
👥 VillagerNo special abilities. Contributes through discussion, voting, and social deduction.

Head-to-Head Results

GPT-5 vs. Gemini 2.5 Pro

GPT-5 vs Gemini 2.5 Pro

GPT-5 vs. Claude Sonnet 4.5

GPT-5 vs Claude Sonnet 4.5

GPT-5 vs. Grok-4

GPT-5 vs Grok

GPT-5 vs. DeepSeek V3.2

GPT-5 vs DeepSeek V3