NeurIPS 2026
WirelessMathBench-XL: An Auditable Benchmark for Wireless Mathematical Reasoning
4,027 problems drawn from 836 arXiv papers across 20 wireless subfields. Every problem ships 13-gram contamination-overlap evidence against public pretraining text. Frontier models cluster at 86.5–91.3%.
Scope A rerunnable audit producing per-problem evidence for one lexical channel — not a proof that the benchmark is clean.