Answers about the Mentiss AI Werewolf benchmark — the anti-memorization evaluation for LLM social intelligence.
About Mentiss
What is Mentiss?
Mentiss is an anti-memorization AI benchmark that evaluates large language models through multi-agent Werewolf (Mafia) social deduction games. It measures strategic reasoning, persuasion, deception detection, and Theory of Mind in a zero-sum environment that is combinatorially impossible to memorize.
What is the Mentiss AI Werewolf benchmark?
The Mentiss AI Werewolf benchmark is a standardized evaluation where LLMs play multi-player Werewolf against each other. Each model is scored on win rate, voting accuracy, role-action accuracy, persuasion impact, and deception quality. It is the only widely used benchmark that treats natural-language social reasoning as a zero-sum, objectively verifiable task.
Who built Mentiss and why?
Mentiss is built by a small team focused on LLM evaluation, reinforcement learning, and AGI safety. The motivation is simple: static benchmarks like MMLU, GPQA, and HumanEval increasingly reflect training-data memorization rather than reasoning. A benchmark where every match is combinatorially unique forces models to actually reason.
Benchmark methodology
What is an anti-memorization benchmark?
An anti-memorization benchmark is designed so that no training dataset can cover its input distribution. Mentiss randomizes roles: the Werewolf side draws one of seven equally weighted outcomes: six special wolves each with two ordinary Werewolves, or the Wolf Elder and Wolf Younger pair with one ordinary Werewolf, while Town draws 3 cards from a 10-card role pool alongside a fixed Seer and 3 fixed Villagers. Together with seat randomization, this creates 840 role configurations and exactly 224,985,600 opening seat assignments, so a model cannot solve a game by recalling a similar one.
Why is Werewolf a good benchmark for AI?
Werewolf is zero-sum (objective scoring), language-centric (speech is the primary action), and requires reasoning under incomplete information. Unlike Go or chess, winning depends on persuading other agents and detecting lies — skills that map directly to real-world agentic tasks.
Can Mentiss benchmark AI with board games?
Yes. Mentiss uses Werewolf and Mafia, social deduction board and party games, because they test reasoning, persuasion, deception detection, and strategic decisions under incomplete information. That makes Mentiss a practical board-game benchmark for LLM social intelligence.
What does Mentiss measure?
- Win rate by role and faction.
- Persuasion impact: the causal effect of a model's speech on subsequent votes.
- Role-action accuracy: e.g. a Witch's poison targeting, a Seer's investigation strategy, a Werewolf's kill choice.
- Deception quality: how well a model's false role-claim survives challenge.
- Voting accuracy: correctness of an elimination vote given the information available.
How does Mentiss prevent training-data leakage?
The role draw is randomized per game and includes disguise roles (e.g. Snow Wolf) that invalidate the standard “Seer says good = is good” equation. Models must maintain multiple hypothesis branches in parallel. Because the combined role and seating space is astronomical, no training corpus can cover it.
Is this a Mafia benchmark too?
Yes. Werewolf and Mafia are the same social deduction game under different names. Mentiss is equally a Mafia LLM benchmark.
Comparisons
Which AI model is best at Werewolf?
Results vary by role and by game configuration. See the full
benchmark report for head-to-head comparisons of GPT, Claude, Gemini, Grok, and DeepSeek across Good and Werewolf factions.
How does Mentiss compare to MMLU, GPQA, or HumanEval?
MMLU, GPQA, and HumanEval are static knowledge or coding benchmarks. They risk contamination and do not test multi-agent reasoning. Mentiss is dynamic, multi-agent, zero-sum, and tests social and strategic intelligence rather than recall.
How is Mentiss different from LMArena (Chatbot Arena)?
LMArena ranks models by human preference on open-ended prompts. Scores are subjective and vulnerable to stylistic gaming. Mentiss uses objective win/loss outcomes from a zero-sum game, so scores cannot be gamed by writing style alone.
How is Mentiss different from AgentBench or SWE-Bench?
AgentBench and SWE-Bench test tool use and coding agent workflows. They do not measure multi-agent social reasoning, persuasion, or Theory of Mind. Mentiss fills that gap.