Code
openEvaluation code, parser, zero-API rebuild scripts, aggregate tables, figures, schemas, datasheet, Croissant and RAI metadata.
NeurIPS 2026 · Evaluations and Datasets Track
Debate movement is not necessarily improvement.
Illustration, not data. Three copies of one model answer an abstract four-option question (A to D), then exchange answers for three debate rounds. The ring is the panel's majority vote; the round track runs from the initial answers (R0) to the final round (R3).
Overview
Paper
Multi-agent large language model (LLM) debate is often evaluated by whether final answers improve, but movement is not necessarily improvement: the same discussion can rescue an initially wrong majority or destroy an initially correct one. Standard final-accuracy evaluations conflate these opposing mechanisms. We introduce an auditable protocol for homogeneous debate on multiple-choice questions (MCQs) that records each run as a transition ledger over collapse, correction, onset, and signed intervention utility. On 6,925 MMLU-Pro debates, the protocol identifies 253 collapses and a parallel correction ledger that changes how interventions should be judged. Replay experiments reveal the central tradeoff: a leave-one-model-out probe-gated freeze prevents 29 collapses but loses 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy. A compact pre-debate 8-probe screen is a triage signal: its unadjusted family-level association with conditional-collapse risk is high (G = 7, Spearman ρ = 0.893, exact two-sided p = 0.0123), but initial-majority accuracy is a close comparator (ρ = 0.821; family partial ρ = 0.767, p = 0.0877), so we do not treat it as calibrated or capability-adjusted prediction. Round-level traces localize many collapses to the first debate round, where early disagreement can precede both harmful cascades and useful recovery. We release replayable schemas, coders, audits, cost cards, and zero-API rebuild scripts so future model–scaffold rows can be compared under the same denominators and signed utility ledger.
Protocol
A homogeneous panel of three agents, copies of one model, answers a multiple-choice question and then debates for three rounds. The panel decision is the majority vote.
Select a cell to see an example trace. Onset is the round at which a correct majority first breaks.
The correct majority first breaks in round 1, so onset is R1.
Under equal weights, corrections minus collapses, divided by the number of debates, equals final minus initial majority accuracy. The ledger splits that change into its two flows.
An intervention is replayed on the same saved debates and scored by the collapses it prevents, the corrections it loses, and their weighted net.
4 argument strengths × 2 social levels (with and without a peer-majority cue). The screen ranks model families for trace logging. It is a triage signal only.
Results · held-out replay
A leave-one-model-out probe-gated freeze, replayed on the 6,525-debate held-out pool, prevents 29 of 251 collapses and loses 108 of 804 corrections.
Illustration. The stream mixes collapses and corrections 251 : 804, and the gate freezes 29 of 251 and 108 of 804, as in the held-out replay. Debates that neither collapse nor correct are not drawn.
−79debates
Collapse prevention alone can recommend the wrong policy, so score interventions by both flows.
Set how much one prevented collapse counts against one lost correction. With correction weight 1 and collapse weight w, net = 29w − 108 debates, and the sign flips at w = 108 / 29 = 3.72.
Net accuracy points = (29w − 108) / 6,525 × 100. Dots: w = 0.5, 1, 2, 4, 8.
Unadjusted family-level association with conditional-collapse risk, G = 7 families, Spearman ρ, on one scale from 0 to 1.
Initial-majority accuracy is a close comparator (ρ = 0.821; family partial ρ = 0.767, p = 0.0877), so the screen is not treated as calibrated or capability-adjusted prediction. It ranks model families for trace logging.
Round-level traces localize many collapses to the first debate round, where early disagreement can precede both harmful cascades and useful recovery.
Release
Three tiers, by how much model-generated text they carry.
Evaluation code, parser, zero-API rebuild scripts, aggregate tables, figures, schemas, datasheet, Croissant and RAI metadata.
Every probe and debate trace with the model-generated text removed. Parsed answers, correctness labels, round structure, token counts and costs are kept: enough to rebuild the paper's tables without API calls.
Full-text traces and the very-strong (convince-wrong) and social-pressure probe templates, which could be reused as a misleading-answer corpus.
No model weights are released.
Cite
@inproceedings{
li2026debateledger,
title={Measuring Collapse and Correction in Homogeneous-Panel {LLM} Debate},
author={Xin Li and Mengbing Liu and Chau Yuen},
booktitle={The Fortieth Annual Conference on Neural Information Processing Systems Evaluations and Datasets Track},
year={2026}
}