Xin Li李鑫

Ph.D. student · Nanyang Technological University · advised by Prof. Chau Yuen

Xin Li李鑫

I work on LLM agents — evaluating what they can do, training them, and orchestrating them with tools, verifiers and each other to get real tasks done. Lately, on agents that improve themselves.

I’m looking for research internships in 2027, on LLM agents, post-training, and self-improvement.

Portrait of Xin Li

The research, as one loopclockwise from 01

01

Evaluate

What can models do, where do they fail, and how reliably can we tell?

02

Train

How can models learn from verifiers, rewards, and other models?

  • DN-MOPD

    In multi-teacher distillation, normalizing each teacher's feedback scale per domain beats label routing alone.

    Preprint 2026

  • Re:Form

    Reinforcement learning for formal software verification in Dafny, with fewer human priors.

    TMLR 2026

  • GraphReduce

    EMNLP 2026 Industry

03

Orchestrate

How should models, tools, and verifiers work together to complete tasks reliably?

  • DebateLedger

    Distinguishes harmful collapse from useful correction in multi-agent debate.

    NeurIPS 2026

  • TLVC

    No single verifier is best for every generator in best-of-K selection; label-free statistics predict which kind wins.

    Findings of EMNLP 2026

  • LACP

    AI4NextG @ NeurIPS 2025

03 → 01. Every new way of orchestrating models creates new failure modes, which need to be evaluated in turn.

So far, mostly where an answer can be checked:formal verificationmathematicscodeNow, self-improvement beyond checkable answers.

RecentAll news →

  1. Our paper SIM-D2NN (on onboard terrain classification straight from raw SAR data, using a stacked metasurface as the classifier) was accepted to IEEE Transactions on Signal Processing.
  2. Two of our papers were accepted to the NeurIPS 2026 Evaluations and Datasets Track: WirelessMathBench-XL, a wireless-math benchmark shipped with a rerunnable contamination audit, and DebateLedger, a protocol separating harmful collapse from useful correction in multi-agent LLM debate.
  3. Two of our papers were accepted to EMNLP 2026: TLVC, on picking the right verifier for best-of-K reasoning selection (Findings), and GraphReduce, on coverage-preserving LLM aggregation of e-commerce reviews (Industry Track).
  4. Our paper RobustMAD (a robustness benchmark for multimodal small language models in anomaly detection) was accepted to TMLR.
  5. Our paper Re:Form (on cutting human priors from RL-trained formal software verification) was accepted to TMLR.

Before the Ph.D.

Robot perception — visual-inertial odometry at MEGVII, RGB-D + IMU indoor mapping at Microsoft Research Asia, and multimodal localization at Gausium Robotics, where I led a five-engineer team and shipped to a fleet of 1,000+ commercial cleaning robots. More →

Contact

Looking for research internships in 2027, and always glad to talk about research or collaboration.