Xin Li李鑫

03 / 03

Orchestrate

How should models, tools, and verifiers work together to complete tasks reliably?

The system around a model at inference time: harnesses and tools, verifier choice, and protocols for several agents working together — including when that system breaks answers that were already right.

Work on this thread3 papers

NeurIPS 2026

Measuring Collapse and Correction in Homogeneous-Panel LLM Debate

Xin Li*, Mengbing Liu*, Chau Yuen

Debate can turn a wrong answer right, and it can also talk a correct majority into a wrong one. DebateLedger counts the two apart, and finds that stopping collapse also stops correction: across 6,925 logged MMLU-Pro debates, 253 collapses turned a correct majority wrong, and a probe-gated freeze prevents 29 of them but gives up 108 corrections — net −79 under equal weights. 58.9% of collapses begin in the first debate round.

Findings of EMNLP 2026

Target-Local Verifier Choice in Best-of-K Reasoning Selection

Xin Li, Hao Jiang, Weisi Lin

On a 34-generator, 7-verifier Best-of-K math panel, one strong process reward model is the best fixed verifier overall — and still not the best verifier for every generator. Label-free candidate statistics predict which verifier class wins.