01
Evaluate
What can models do, where do they fail, and how reliably can we tell?
- WirelessMathBench-XL
A wireless-math benchmark with per-problem provenance and reproducible text-overlap audits.
NeurIPS 2026
- DafnyComp
Models that verify functions one at a time can still fail once the specifications have to compose.
ICLR 2026
- LiveCANNBench
Findings of ACL 2026
- RobustMAD
TMLR 2026
- WirelessMathBench
Findings of ACL 2025