Xin Li李鑫

02 / 03

Train

What should a model be rewarded for — especially when there is no answer key?

What the training signal should be — which reward, which verifier, and how much each teacher's feedback should count.

Work on this thread4 papers

Findings of EMNLP 2026

Target-Local Verifier Choice in Best-of-K Reasoning Selection

Xin Li, Hao Jiang, Weisi Lin

On a 34-generator, 7-verifier Best-of-K math panel, one strong process reward model is the best fixed verifier overall — and still not the best verifier for every generator. Label-free candidate statistics predict which verifier class wins.

EMNLP 2026 Industry

GraphReduce: Coverage-Preserving LLM Aggregation for E-commerce Review Insights

Hao Jiang, Xin Li, Yichi Zhang, Weisi Lin

Turns thousands of review tuples into a ranked list of product insights while every tuple stays attached to the insight it supports: membership is computed outside the LLM, so coverage is 100% with no duplicates. An anti-copy reinforcement objective lifts the upstream extractor from 59.3% to 93.1% production quality.

Also on this thread