QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents

Published in arXiv preprint, 2026

Status: preprint on arXiv · My role: contributing author.

Paper (arXiv) Code

Summary

Social deduction games are a popular way to probe reasoning, deception, coordination, and belief modeling in LLM agents — but most environments score only game outcomes such as win rate, and most are text-only. That makes it impossible to tell whether an agent’s language is actually grounded in what it perceived and did.

QUACK is an open-source environment and evaluation framework that audits grounding directly. Agents navigate configurable graph-based maps under partial observability, see rendered global and local views, complete location-bound tasks, discuss freely, and vote under hidden-role adversarial incentives.

The core Statement Verification Pipeline reconstructs each agent’s ground-truth trajectory from engine logs and checks every discussion claim against it, automatically flagging:

  • spatial hallucination — claims about places the agent never observed
  • unsupported accusation — accusations with no grounded evidence
  • deception collapse and language–action inconsistency

Finding: across three frontier VLMs in both homogeneous and cross-model adversarial settings, even the strongest agent hallucinates 15.1% of its verifiable spatial claims and makes over half of its accusations without grounded evidence.

Work with the Mila group led by Ye Yuan and Prof. Xue (Steve) Liu.

Recommended citation: Ye Yuan, Rui Song, Weien Li, et al. (including Yonghan Yang), and Xue Liu. (2026). "QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents." arXiv:2605.27068.
Download Paper