01
Measure
Does a score mean what it appears to mean — and would we notice if it didn't?
Ph.D. student · Nanyang Technological University · advised by Prof. Chau Yuen
I work on LLM agents — measuring what they can do, training them, and finding out what happens when several work together. Lately, on agents that improve themselves.
01
Does a score mean what it appears to mean — and would we notice if it didn't?
02
What should a model be rewarded for — especially when there is no answer key?
03
When several models work together, does the interaction help — or quietly destroy answers that were already right?
03 → 01. Coordination creates new failure modes, which have to be measured too.
So far, mostly where an answer can be checked:formal verificationmathematicscodeNow, self-improvement where it can't.
Robot perception — visual-inertial odometry at MEGVII, RGB-D + IMU indoor mapping at Microsoft Research Asia, and multimodal localization at Gausium Robotics, where I led a five-engineer team and shipped to a fleet of 1,000+ commercial cleaning robots. More →
Always glad to talk about research or collaboration.