About
I am a Ph.D. student at Nanyang Technological University (NTU), advised by Prof. Chau Yuen. I work on making large language models (LLMs) reason reliably: benchmarks that measure whether they do, post-training that teaches them to, and agentic systems that test whether coordination actually helps. Most of it lands in domains where an answer can be checked—formal verification, mathematics, and code.
- LLM Evaluation Benchmarks and evaluation protocols built to be audited rather than trusted: per-problem provenance, verifier-checkable answers, and explicit limits on what a score supports—including WirelessMathBench-XL, DafnyComp, WritingPreferenceBench, and RobustMAD.
- Post-training Reinforcement learning, on-policy distillation, and verifier-guided selection, all aimed at reasoning where the reward can be checked rather than guessed—including Re:Form, WirelessMathLM, TLVC, ListOPD, and DN-MOPD.
- Agentic Systems Protocols, memory, and measurement for systems where several models interact: what they send each other, what they keep, and whether the interaction helps or hurts—including LACP and DebateLedger.
Before my Ph.D. I worked on robot perception—visual-inertial odometry at MEGVII, RGB-D + IMU indoor mapping at Microsoft Research Asia, and multimodal localization at Gausium Robotics, where I led a five-engineer team and shipped to a fleet of 1,000+ commercial cleaning robots. I received my M.E. from Peking University and my B.E. from Northeastern University, China.
I'm always glad to talk about research or collaboration—just email me at xin019@e.ntu.edu.sg.
News
- Sep 2026 Our paper SIM-D2NN (on onboard terrain classification straight from raw SAR data, using a stacked metasurface as the classifier) was accepted to IEEE Transactions on Signal Processing.
- Sep 2026 Two of our papers were accepted to the NeurIPS 2026 Evaluations and Datasets Track: WirelessMathBench-XL, a wireless-math benchmark shipped with a rerunnable contamination audit, and DebateLedger, a protocol separating harmful collapse from useful correction in multi-agent LLM debate.
- Aug 2026 Two of our papers were accepted to EMNLP 2026: TLVC, on picking the right verifier for best-of-K reasoning selection (Findings), and GraphReduce, on coverage-preserving LLM aggregation of e-commerce reviews (Industry Track).
- Jun 2026 Our paper RobustMAD (a robustness benchmark for multimodal small language models in anomaly detection) was accepted to TMLR.
- May 2026 Our paper Re:Form (on cutting human priors from RL-trained formal software verification) was accepted to TMLR.
- Apr 2026 Our paper LiveCANNBench (a benchmark for AI coding on Ascend CANN) was accepted to Findings of ACL 2026.
- Jan 2026 Received a Google Gemini Academic Program Award (US$10,000).
- Jan 2026 Our paper DafnyComp (on benchmarking LLMs for compositional formal verification) was accepted to ICLR 2026.
- Dec 2025 Received the Rohde & Schwarz Award at the IEEE 6G Summit Singapore.
- Sep 2025 Our paper LACP (a communication protocol for LLM agents) was accepted to AI4NextG @ NeurIPS 2025.
- May 2025 Our paper WirelessMathBench (a benchmark for mathematical reasoning in wireless communications) was accepted to Findings of ACL 2025.
- May 2025 Our workshop on Advancements for Intelligent Robotics in 4D Scenes at IROS 2025 was accepted.
- Mar 2025 Our paper SIM-D2NN (on onboard terrain classification for remote sensing) was accepted to the ML4RS @ ICLR 2025.
- Jan 2025 Our paper TransPathNet (on indoor pathloss prediction) was accepted to ICASSP 2025, ranked 4th in the Indoor Pathloss Challenge.
- Jan 2025 Started my Ph.D. at NTU, supported by the NTU Research Scholarship.
Publications
* equal contribution · full list → Google Scholar
-
GraphReduce: Coverage-Preserving LLM Aggregation for E-commerce Review Insights
EMNLP Industry Track · 2026
-
LACP: LLM Agent Communication Protocol Requires Urgent Standardization
AI4NextG @ NeurIPS · 2025 -
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
Findings of ACL · 2025
Experience
-
Research Assistant — NTU Singapore
Apr 2024 – Jan 2025
Supervised by Prof. Chau Yuen, IEEE Fellow. Built WirelessMathBench (Findings of ACL 2025), the first benchmark for LLM mathematical reasoning in wireless communications, and TransPathNet (ICASSP 2025), which placed 4th in the Indoor Pathloss Prediction Challenge at 9.73 dB RMSE.
-
SLAM Algorithm Engineer (Project Lead) — Gausium Robotics, Singapore
Mar 2022 – Feb 2024
Led a team of five engineers building hierarchical multimodal localization (vision, LiDAR, and Wi-Fi) for commercial cleaning robots: 95% localization accuracy across 100,000+ m² of complex environments; the system is deployed on a global fleet of 1,000+ active robots.
-
Research Intern — Microsoft Research Asia, Beijing
Sep 2020 – Mar 2021
Supervised by Dr. Yang Liu and Dr. Yizhong Zhang. Built a multi-sensor fusion system (RGB-D + IMU) for large-scale indoor mapping, producing vectorized maps of 10,000+ m² of commercial space at sub-meter accuracy.
-
Research Intern — MEGVII, Beijing
Feb 2019 – Mar 2020
Supervised by Dr. Yijia He. Developed real-time monocular visual-inertial odometry from heterogeneous features, reaching semi-dense 3D mesh reconstruction at 30+ FPS; the work led to a first-author paper at IROS 2020.
Mentoring
-
Chengqi Liang — M.Sc. Dissertation, Nanyang Technological University
2026
Went on to a Ph.D. at The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen).
-
Yukun Jin — B.Eng. Final Year Project, Nanyang Technological University
2026
Undergraduate at Wuhan University under NTU's 3.5+0.5+1 integrated programme; went on to an M.Sc. at NTU.
-
Haoyu Xu — M.Comp. Dissertation, National University of Singapore
2024
Went on to a Ph.D. at Peking University.
Education
-
Ph.D. — Nanyang Technological University, Singapore
2025 – 2029 (expected)
Supervised by Prof. Chau Yuen, IEEE Fellow.
-
M.E. — Peking University, China
2018 – 2021
Supervised by Prof. Jinlong Lin.
- B.E. — Northeastern University, China 2014 – 2018
Awards
Research Grants
- Google Gemini Academic Program Award (US$10,000)2026
- Modal Academics Compute Grant (US$2,000)2025
- Cohere Labs Catalyst Grant (US$1,500)2025
- OpenAI Researcher Access Program (US$1,000)2025
Academic Honors
- Rohde & Schwarz Award (IEEE 6G Summit Singapore)2025
- PREMIA Best Student Paper Award Finalist2025
- NTU Research Scholarship (Full Ph.D. Funding)2025
Talks
-
Teaching Large Language Models Mathematical Reasoning in Wireless Communications: From Benchmarking to Efficient Training
October 27, 2025 -
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
June 11, 2025
Professional Service
- Conference Reviewer NeurIPS, ICLR, ICML, AAAI, CVPR, ECCV, AISTATS, SIGGRAPH, IROS, ICRA.
- Journal Reviewer IEEE RA-L, ACM TOG, IEEE TNNLS.
- Workshop Organizer AIR4D@IROS 2025.