TL;DR: RLVR has proven transformative for math and code—domains where verification is binary. Social intelligence is different: actions are probabilistically correct, outcomes depend on other agents, and games unfold over multiple rounds.
Mentiss introduces a fixed, prompt-engineered "expert model" (the Oracle Agent, with a god's-eye view) to assist verification: it is an external judge model that does not participate in the game and knows the full ground truth. Given the complete game state, it produces a stable coherence score, helping offset the sparsity and noise of outcome-only rewards.
Mentiss addresses the RLVR challenge through three mechanisms:
- GRPO: Verify the Unverifiables — Using game outcomes to score rhetoric
- Distributional Verification — Evaluating probability distributions, not single outputs
- Phase-Dependent Rewards — Verification that adapts to information availability
The Setup: Werewolf (Mafia)
In Werewolf, hidden Werewolves try to eliminate Villagers while the Good faction tries to identify and vote them out. Each round has two phases:
- Night: Special roles act (e.g., the Witch can poison one player)
- Day: Players give speeches, then vote to eliminate someone
The Witch is a Good role who knows nothing about other players' identities but can observe speeches and voting patterns. We use her as our running example.
1. GRPO: Verify the Unverifiables
Persuasion is subjective—what makes a speech convincing? We use GRPO (Group Relative Policy Optimization): generate multiple candidate responses, let the game evaluate them, reward based on relative performance.
For GRPO, we decompose the reward into two parts and take a weighted sum:
- External outcome: the voting result (did you successfully lead the Good faction to eliminate the fake Witch?)
- Internal coherence: the Oracle Agent's coherence score (is it consistent across turns, and consistent with the facts and rules so far?)
Example: Counter-Claim
A Werewolf falsely claims to be the Witch. The true Witch speaks last.
We generate 64 candidate counter-speeches. After voting:
- 4/5 Villagers vote to eliminate the fake → High score
- Vote split 50-50 → Medium score
- Villagers vote against the true Witch → Failed
No human annotator needed—the voting outcome defines "good rhetoric," and the Oracle Agent's coherence score penalizes obvious contradictions or fabricated information, making the training signal more stable.
2. Distributional Verification
Instead of "who do you poison?", we ask "rank all players by your inclination to poison them."
For Distributional Verification, we similarly decompose the reward into two parts and take a weighted sum:
- Distributional hit score: the more accurately you "guess" (i.e., the higher probability mass you assign to true Werewolves), the higher the score
- Internal coherence: the Oracle Agent's coherence score (is your probability assignment consistent with your stated reasoning and the game history?)
Intuitively, we can also apply a variance / concentration-style shaping term to the probability distribution: reward placing probability mass more sharply on a small number of the most suspicious targets, and penalize overly uniform "spray-and-pray" allocations—this makes the signal sharper and easier to learn from.
Example: Ranked Targets
7 players remain, 3 are Werewolves. The Witch outputs:
- Player 4 — 0.35
- Player 2 — 0.28
- Player 6 — 0.20
- Player 1 — 0.10
- Player 5 — 0.05
- Player 7 — 0.02
Ground Truth: Players 4, 2, and 7 are Werewolves.
- Top 3 contains 2/3 Werewolves → High reward
- Top 3 with all 3 Werewolves → Maximum reward
Gradient-rich signals instead of binary pass/fail.
3. Phase-Dependent Rewards
The same action can be rational or irrational depending on when it occurs.
Example: Timing Matters
The Witch poisons Player 4 (a Villager who gave suspicious speeches).
Night 2: Information is scarce → ✅ Soft Pass (good process under uncertainty)
Night 4: Voting history should have revealed the truth → ❌ Hard Fail
Reward Multipliers:
- Early game: 2.0x
- Mid game: 1.0x
- Late game: 0.5x
The Bigger Picture
This was never about mastering a social deduction game. It's about a roadmap to AGI
It's about LLMs learning to persuade, to read between the lines, to model other minds. Rhetoric is one of the missing pieces toward AGI. Math and code gave us verifiable reasoning. Social games can give us verifiable rhetoric.
We don't claim to have solved social intelligence. But Mentiss can be one small contribution toward building AI that truly understands how humans think, communicate, and cooperate.
References: