NeurIPS 2026 · Evaluations and Datasets Track

Measuring Collapse and Correction in Homogeneous-Panel LLM Debate

Debate movement is not necessarily improvement.

Xin Li*, Mengbing Liu*, Chau YuenNanyang Technological University · *Equal contribution

Illustration A correct majority collapses correct option B

Illustration, not data. Three copies of one model answer an abstract four-option question (A to D), then exchange answers for three debate rounds. The ring is the panel's majority vote; the round track runs from the initial answers (R0) to the final round (R3).

6,925 MMLU-Pro debates in the trace set
253 collapses in that set
29vs108 collapses prevented vs corrections lost by a probe-gated freeze (held-out replay)
−1.21 accuracy points, net under equal weights

Overview

The same debate can rescue a majority or break one

Debate movement is not necessarily improvement. The same scaffold can collapse an initially correct majority or correct an initially wrong one, and signed replay scores an intervention by both the collapses it prevents and the corrections it loses.

Paper

Abstract

Multi-agent large language model (LLM) debate is often evaluated by whether final answers improve, but movement is not necessarily improvement: the same discussion can rescue an initially wrong majority or destroy an initially correct one. Standard final-accuracy evaluations conflate these opposing mechanisms. We introduce an auditable protocol for homogeneous debate on multiple-choice questions (MCQs) that records each run as a transition ledger over collapse, correction, onset, and signed intervention utility. On 6,925 MMLU-Pro debates, the protocol identifies 253 collapses and a parallel correction ledger that changes how interventions should be judged. Replay experiments reveal the central tradeoff: a leave-one-model-out probe-gated freeze prevents 29 collapses but loses 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy. A compact pre-debate 8-probe screen is a triage signal: its unadjusted family-level association with conditional-collapse risk is high (G = 7, Spearman ρ = 0.893, exact two-sided p = 0.0123), but initial-majority accuracy is a close comparator (ρ = 0.821; family partial ρ = 0.767, p = 0.0877), so we do not treat it as calibrated or capability-adjusted prediction. Round-level traces localize many collapses to the first debate round, where early disagreement can precede both harmful cascades and useful recovery. We release replayable schemas, coders, audits, cost cards, and zero-API rebuild scripts so future model–scaffold rows can be compared under the same denominators and signed utility ledger.

Protocol

Record every run as a transition ledger

A homogeneous panel of three agents, copies of one model, answers a multiple-choice question and then debates for three rounds. The panel decision is the majority vote.

Setting

A homogeneous panel

  • 3agents, all copies of one model
  • 3debate rounds after the initial answers
  • 1panel decision: the majority vote
Transition ledger

Initial majority crossed with final majority

Select a cell to see an example trace. Onset is the round at which a correct majority first breaks.

Final correct Final wrong Initial correct Initial wrong
Example majority trace (illustration)
  1. R0✓
  2. R1✗
  3. R2✗
  4. R3✗

The correct majority first breaks in round 1, so onset is R1.

Accounting identity

The ledger decomposes accuracy change

corrections − collapsesdebates = final acc − initial acc

Under equal weights, corrections minus collapses, divided by the number of debates, equals final minus initial majority accuracy. The ledger splits that change into its two flows.

Signed utility

Score an intervention by both flows

net= w · collapses prevented−1 · corrections lost

An intervention is replayed on the same saved debates and scored by the collapses it prevents, the corrections it loses, and their weighted net.

Pre-debate screen

Eight probes before any debate

4 argument strengths × 2 social levels (with and without a peer-majority cue). The screen ranks model families for trace logging. It is a triage signal only.

social level weak moderate strong very strong
no peer cue 1 2 3 4
peer-majority cue 5 6 7 8
Eight-probe instrument and triage workflow. The total flip rate αtot ranks model families for trace logging; splitting by initial correctness gives adversarial and corrective flip rates.

Results · held-out replay

The probe-gated freeze prevents 29 collapses and loses 108 corrections

A leave-one-model-out probe-gated freeze, replayed on the 6,525-debate held-out pool, prevents 29 of 251 collapses and loses 108 of 804 corrections.

Illustration. The stream mixes collapses and corrections 251 : 804, and the gate freezes 29 of 251 and 108 of 804, as in the held-out replay. Debates that neither collapse nor correct are not drawn.

Drawn to scale · one count axis
Collapses prevented29 of 251 · 11.6%
Corrections lost108 of 804 · 13.4%
Bar length = debates. Pale: all collapses or corrections in the pool. Solid: frozen by the gate.0 to 804
Net under equal weights

−79debates

Net accuracy change
−1.21 points
Held-out pool
6,525 debates
Break-even collapse weight
108 / 29 = 3.72

Collapse prevention alone can recommend the wrong policy, so score interventions by both flows.

Signed utility · interactive

Weigh the two flows

Set how much one prevented collapse counts against one lost correction. With correction weight 1 and collapse weight w, net = 29w − 108 debates, and the sign flips at w = 108 / 29 = 3.72.

net = 29 × 1.00 − 108 = −79 debates −1.21 points Weighted net on 6,525 held-out debates. Below break-even the freeze scores negative.

Net accuracy points = (29w − 108) / 6,525 × 100. Dots: w = 0.5, 1, 2, 4, 8.

Pre-debate screen · family level

The screen is a triage signal

Unadjusted family-level association with conditional-collapse risk, G = 7 families, Spearman ρ, on one scale from 0 to 1.

8-probe screenρ = 0.893 · exact two-sided p = 0.0123
Initial-majority accuracy (comparator)ρ = 0.821

Initial-majority accuracy is a close comparator (ρ = 0.821; family partial ρ = 0.767, p = 0.0877), so the screen is not treated as calibrated or capability-adjusted prediction. It ranks model families for trace logging.

Round-level traces

Where collapses start

Round-level traces localize many collapses to the first debate round, where early disagreement can precede both harmful cascades and useful recovery.

Release

Rebuild the tables without API calls

Three tiers, by how much model-generated text they carry.

Code

open
GitHub · LiXin97/DebateLedger

Evaluation code, parser, zero-API rebuild scripts, aggregate tables, figures, schemas, datasheet, Croissant and RAI metadata.

MIT for codeCC BY 4.0 for tables, figures, schemas

Open data

open
Hugging Face · XINLI1997/DebateLedger

Every probe and debate trace with the model-generated text removed. Parsed answers, correctness labels, round structure, token counts and costs are kept: enough to rebuild the paper's tables without API calls.

CC BY 4.0

Gated data

on request
Hugging Face · XINLI1997/DebateLedger-gated

Full-text traces and the very-strong (convince-wrong) and social-pressure probe templates, which could be reused as a misleading-answer corpus.

Research-use licenseRequests reviewed by the authors

No model weights are released.

Cite

BibTeX

li2026debateledger
@inproceedings{
li2026debateledger,
title={Measuring Collapse and Correction in Homogeneous-Panel {LLM} Debate},
author={Xin Li and Mengbing Liu and Chau Yuen},
booktitle={The Fortieth Annual Conference on Neural Information Processing Systems Evaluations and Datasets Track},
year={2026}
}
Xin Li and Mengbing Liu contributed equally. lixin.ai