← Back to all articles

Extending RLVR Beyond Math and Code: Verifiable Rewards for Rhetoric

This isn't about LLMs mastering a social deduction game. It's about LLMs mastering rhetoric - the ability to persuade, deceive, and reason about other minds. We believe this is one piece of the AGI puzzle, and Mentiss's mechanisms here can be a small contribution toward that goal.

Extending RLVR Beyond Math and Code: Verifiable Rewards for Rhetoric

TL;DR: RLVR has proven transformative for math and code—domains where verification is binary. Social intelligence is different: actions are probabilistically correct, outcomes depend on other agents, and games unfold over multiple rounds.


Mentiss introduces a fixed, prompt-engineered "expert model" (the Oracle Agent, with a god's-eye view) to assist verification: it is an external judge model that does not participate in the game and knows the full ground truth. Given the complete game state, it produces a stable coherence score, helping offset the sparsity and noise of outcome-only rewards.

Mentiss addresses the RLVR challenge through three mechanisms:

  1. GRPO: Verify the Unverifiables — Using game outcomes to score rhetoric
  2. Distributional Verification — Evaluating probability distributions, not single outputs
  3. Phase-Dependent Rewards — Verification that adapts to information availability

The Setup: Werewolf (Mafia)

In Werewolf, hidden Werewolves try to eliminate Villagers while the Good faction tries to identify and vote them out. Each round has two phases:

  • Night: Special roles act (e.g., the Witch can poison one player)
  • Day: Players give speeches, then vote to eliminate someone

The Witch is a Good role who knows nothing about other players' identities but can observe speeches and voting patterns. We use her as our running example.


1. GRPO: Verify the Unverifiables

Persuasion is subjective—what makes a speech convincing? We use GRPO (Group Relative Policy Optimization): generate multiple candidate responses, let the game evaluate them, reward based on relative performance.

For GRPO, we decompose the reward into two parts and take a weighted sum:

  • External outcome: the voting result (did you successfully lead the Good faction to eliminate the fake Witch?)
  • Internal coherence: the Oracle Agent's coherence score (is it consistent across turns, and consistent with the facts and rules so far?)

Example: Counter-Claim

A Werewolf falsely claims to be the Witch. The true Witch speaks last.

We generate 64 candidate counter-speeches. After voting:

  • 4/5 Villagers vote to eliminate the fake → High score
  • Vote split 50-50 → Medium score
  • Villagers vote against the true Witch → Failed

No human annotator needed—the voting outcome defines "good rhetoric," and the Oracle Agent's coherence score penalizes obvious contradictions or fabricated information, making the training signal more stable.


2. Distributional Verification

Instead of "who do you poison?", we ask "rank all players by your inclination to poison them."

For Distributional Verification, we similarly decompose the reward into two parts and take a weighted sum:

  • Distributional hit score: the more accurately you "guess" (i.e., the higher probability mass you assign to true Werewolves), the higher the score
  • Internal coherence: the Oracle Agent's coherence score (is your probability assignment consistent with your stated reasoning and the game history?)

Intuitively, we can also apply a variance / concentration-style shaping term to the probability distribution: reward placing probability mass more sharply on a small number of the most suspicious targets, and penalize overly uniform "spray-and-pray" allocations—this makes the signal sharper and easier to learn from.

Example: Ranked Targets

7 players remain, 3 are Werewolves. The Witch outputs:

  1. Player 4 — 0.35
  2. Player 2 — 0.28
  3. Player 6 — 0.20
  4. Player 1 — 0.10
  5. Player 5 — 0.05
  6. Player 7 — 0.02

Ground Truth: Players 4, 2, and 7 are Werewolves.

  • Top 3 contains 2/3 Werewolves → High reward
  • Top 3 with all 3 Werewolves → Maximum reward

Gradient-rich signals instead of binary pass/fail.


3. Phase-Dependent Rewards

The same action can be rational or irrational depending on when it occurs.

Example: Timing Matters

The Witch poisons Player 4 (a Villager who gave suspicious speeches).

Night 2: Information is scarce → ✅ Soft Pass (good process under uncertainty)

Night 4: Voting history should have revealed the truth → ❌ Hard Fail

Reward Multipliers:

  • Early game: 2.0x
  • Mid game: 1.0x
  • Late game: 0.5x

The Bigger Picture

This was never about mastering a social deduction game. It's about a roadmap to AGI

It's about LLMs learning to persuade, to read between the lines, to model other minds. Rhetoric is one of the missing pieces toward AGI. Math and code gave us verifiable reasoning. Social games can give us verifiable rhetoric.

We don't claim to have solved social intelligence. But Mentiss can be one small contribution toward building AI that truly understands how humans think, communicate, and cooperate.


References: