NeurIPS 2026
WirelessMathBench-XL: An Auditable Benchmark for Wireless Mathematical Reasoning
A benchmark that ships its own audit: every problem carries 13-gram contamination-overlap evidence against public pretraining text. 4,027 problems drawn from 836 arXiv papers across 20 wireless subfields; frontier models cluster at 86.5–91.3%.
Scope A rerunnable audit producing per-problem evidence for one lexical channel — not a proof that the benchmark is clean.