Xin Li李鑫

01 / 03

Measure

Does a score mean what it appears to mean — and would we notice if it didn't?

Benchmarks and evaluation protocols built to be audited rather than trusted: per-problem provenance, verifier-checkable answers, and an explicit limit on what a score supports.

Work on this thread5 papers

NeurIPS 2026

WirelessMathBench-XL: An Auditable Benchmark for Wireless Mathematical Reasoning

Xin Li, Mengbing Liu, Yiyang Zhu, Wenhe Zhang, Li Wei, Jiancheng An, Chau Yuen

4,027 problems drawn from 836 arXiv papers across 20 wireless subfields. Every problem ships 13-gram contamination-overlap evidence against public pretraining text. Frontier models cluster at 86.5–91.3%.

Scope A rerunnable audit producing per-problem evidence for one lexical channel — not a proof that the benchmark is clean.

Findings of ACL 2026

LiveCANNBench: Benchmark SWE AI Coding for Ascend CANN

Sijie Wang, Kai Zhao, Wee Peng Tay, Shuo Zhang, Chengwen Liu, Quanjiang Guo, Ren Junhao, Xin Li, Heng Lian, Jingdi Lei, Rui She, Huacan Wang, Ronghao Chen

400+ SWE-level task instances built from real Ascend CANN repositories — multi-file, multi-language and execution-aware — on a live benchmarking paradigm that mitigates leakage.

Also on this thread