WMB-XL
A base station's antenna array radiating waves that turn into mathematics A lattice mast topped by a 4-by-4 antenna array sends out circular waves and two beams. One beam lands on a 16-QAM constellation diagram; the other becomes the capacity formula log2(1 + SNR). Between them float a summation sign, a 2-by-2 channel matrix H, the Hermitian product H-superscript-H H, and the positive-part operator. 16-QAM · I/Q I Q Σ h11 h12 h21 h22 H = CAPACITY · bits/s/Hz C = log2(1 + SNR) HHH (·)+ h[t] γ
NeurIPS 2026 Evaluations & Datasets Track

WirelessMathBench-XL

An Auditable Benchmark for Wireless Mathematical Reasoning

  • Xin Li
  • Mengbing Liu
  • Yiyang Zhu
  • Wenhe Zhang
  • Li Wei
  • Jiancheng An
  • Chau Yuen

Nanyang Technological University

4,027problems, verifier-facing ground truths
836retained arXiv source papers, 20 subfields
3task formats: MCQ · Fill-in · Full Equation Completion
12.9 BRedPajama-arXiv 13-grams streamed by the audit
95.7%in the strict zero-hit view S0 (3,853 problems; one audit channel)
TL;DR

WirelessMathBench-XL is a 4,027-problem benchmark for wireless mathematical reasoning, built from 836 arXiv papers. Its distinguishing feature is a reverse-probe audit at a fixed 13-gram threshold: every problem ships with its source-paper ID and per-problem prompt-surface lexical-overlap metadata against RedPajama-arXiv, so reviewers and users can inspect and recompute it. It is an auditable protocol for one channel, not a cleanliness certificate.

Chapter 1 · The problem

Deriving the equation, or remembering the paper?

When a frontier model solves a wireless-math item, it may be deriving the equation from the stated constraints, or it may have seen the source-paper text during pretraining. Two facts make this question sharp.

Wireless math is full of rules nobody writes down

Derivation-grounded wireless problems carry implicit constraints. The load-bearing content of a problem is often the constraint, not only the equation. Four recurring classes describe what the task formats probe:

  • Water-filling: four channels with floors of different heights under a common water level. Three channels get positive power; the fourth floor sits above the water, so its would-be negative power is clipped to zero. water level μ p1 p2 p3 p4 = 0 (μ − floor) < 0
    C1

    Power non-negativity

    Power can’t be poured below zero: when a channel is too weak, its share is clipped to 0, not made negative.

    KKT closed forms can admit negative roots that need a water-filling $(\cdot)^{+}$ never stated in the text.

  • Matrix shapes: a tall N-by-M channel matrix H. The product H-Hermitian times H is a small M-by-M square, while H times H-Hermitian is a large N-by-N square; silently swapping them changes the dimensions. H N×M HHH M×M ≠ HHH N×N silent swap ✗
    C2

    Complex-matrix dimensionality

    Hᴴ H and H Hᴴ are different sizes: swapping them quietly changes the shape of the answer.

    Silent $\mathbf{H}^{H}\mathbf{H}\leftrightarrow\mathbf{H}\mathbf{H}^{H}$ swaps under Hermitian transpose, Kronecker product, inversion.

  • The rate curve log2(1 + gamma). For gamma at least zero the curve rises smoothly; the region gamma below zero is shaded as not a valid SINR, and a point evaluated there is crossed out. γ < 0 invalid valid: γ ≥ 0 log2(1+γ) γ 0
    C3

    Log-domain validity

    A rate log(1+γ) only makes sense for a genuine SINR γ ≥ 0; algebra can push it out of that domain.

    $\log(1+\gamma)$ used outside its domain after aggressive interference cancellation.

  • A timeline of channel samples at t minus 2, t minus 1, t and t plus 1. Arrows from past and present samples feed the decision W at time t; the arrow from the future sample h at t plus 1 is crossed out. future t−2 t−1 t t+1 W[t] h[t+1]
    C4

    Signal causality

    A decision made at time t can use channel samples up to t, never the future h[t+1].

    Filters or precoders conditioning on $\mathbf{h}[t+1]$ inside a $t$-time decision.

These four classes describe what the task formats probe. In v1.0 they are descriptive (not per-instance labels), so they are not used to claim class-level prevalence. The mini-diagrams are explanatory sketches, not benchmark items.

…and the source papers may already be in the training data

Benchmarks built from arXiv papers can overlap the same public text used in LLM pretraining: public resources such as RedPajama, Dolma, and ProofPile-2 include arXiv-derived slices, and exposure varies by model cutoff and training mixture.

For a released technical-domain benchmark, this turns the validity question into an artifact question: does the release give reviewers and future users enough provenance and metadata to re-check the audit?

WirelessMathBench-XL's answer: keep every problem's paper_id, and attach per-problem, recomputable overlap evidence from a corpus-scale audit (Chapter 3), with its limits stated up front (Chapter 4).
How overlap can arise One arXiv paper feeds two paths: it becomes a benchmark item, and it may also appear in a pretraining corpus that an LLM learns from. When the LLM answers the benchmark item, it is unclear whether it derived or recalled the answer. arXiv source paper Benchmark item R = MASK Pretraining corpus arXiv-derived slices LLM derived… or recalled?

Where the 4,027 problems come from

Papers are filtered for wireless relevance and mathematical content, structured records are extracted from each paper's LaTeX, three verifier-facing formats are generated, and every problem passes a two-tier QA (GPT-4o rubric, then expert validation) before audit metadata is attached.

  • ~47,000candidate papers from 24 arXiv categories, July 2005 – August 2025 (deterministic relevance score)
  • 3,186with substantial mathematical content (GPT-4o verifier)
  • 970paper source pool after ranking on rigour, citation signal and topical diversity
  • 836retained papers contribute accepted problems
  • 4,027accepted problems: 4,027 / 5,205 expert-screened candidates (77.4%); problem-level 3,227 / 800 train/test split
Three-panel construction pipeline. Panel 1, Paper Collection: arXiv multi-category crawling, an initial crawl of about 47,000 papers, an LLM relevance filter to 3,186 papers, and manual refinement to a final selection of 970. Panel 2, Formula Extraction with DeepSeek-R1: 10 to 25 formulas per paper turned into MCQ, Fill-in and FEC candidates. Panel 3, Quality Assurance: an automated rubric and five expert reviewers with a quality threshold of 3 out of 5; 78 percent pass and 22 percent fail.
Figure 1 (paper). WirelessMathBench-XL construction pipeline. Counts inside the MCQ, Fill-in, and FEC boxes denote pre-QA generated candidates; the final release contains 4,027 accepted problems.

Chapter 2 · Try it yourself

Three formats, three kinds of hard

Each structured record yields problems in three formats with deliberately different difficulty profiles. Below is a pool of 29 real, publicly released records (CC BY 4.0), shown verbatim. Filter by format or by audit metadata, draw another problem, answer the multiple-choice items, then climb the masking ladder to watch one equation lose its scaffolding.

Format
Audit metadata

MCQ

Multiple choice · record #16130

Distractors are built to violate an implicit constraint

Background

In the UAV-GBS communication link, the real-time data rate depends on the elevation angle and distance. Here, $R_{g_m}(t)$ is the communication rate from the UAV to the $m$-th GBS at time $t$ (in bps), $\theta_{g_m}(t)$ is the elevation angle (in degrees), $d_{g_m}(t)$ is the distance (in meters), $P(t)$ is the UAV's transmission power (in watts), $H$ is the bandwidth (in Hz). The model uses environment-dependent parameters $\chi_1, \chi_2, \chi_3, \chi_4$ (dimensionless, with $\chi_1<0, \chi_2>0, \chi_4>0, \chi_3+\chi_4=1$), a normalized SNR parameter $\hat{\gamma}$ (dimensionless), and the path-loss exponent $\alpha$ (dimensionless).

Question

Which expression completes the capacity component of the communication rate equation?

$R_{g_m}(t) = \left( \chi_3 + \frac{\chi_4}{1 + e^{-(\chi_1 + \chi_2 \theta_{g_m}(t))}} \right) \boxed{\,\color{#b3401c}{?}\,}$
Reveal answer

Ground truth: option B

$R_{g_m}(t) = \left( \chi_3 + \frac{\chi_4}{1 + e^{-(\chi_1 + \chi_2 \theta_{g_m}(t))}} \right) \boxed{\color{#00706c}{H \log_2 \left(1 + \frac{ \hat{\gamma} P(t) }{ (d_{g_m}(t))^{\alpha} } \right)}}$

MCQ distractors are constraint-violating by construction (for example matrix-dimension mismatches, operator-order errors, or sign violations), so the format probes implicit constraints rather than surface plausibility.

Derived from arXiv:2411.02757v1record #16130type MCQtest splitCC BY 4.0RedPajama-arXiv 13-gram hits: 0 · in S0

Mask ladder

One equation, four masking levels

Records #9721–#9724 share one background and one channel-model equation. Each masks a larger share of its components, ending with Full Equation Completion. Same physics, less scaffolding.

Shared background (all four rungs)

In a narrowband MISO channel model for a mmWave/THz system with a Uniform Planar Array (UPA), the channel matrix is represented as a sum of contributions from $L$ propagation paths. The channel matrix $\boldsymbol{H} \in \mathbb{C}^{N \times N}$ represents the wireless channel between the transmitter's $N$ antennas and a single-antenna receiver, where $N$ is the total number of antenna elements in the UPA. $\beta_\ell \in \mathbb{C}$ is the complex gain of the $\ell$-th path (dimensionless), $\omega_{a,\ell} \in [-\pi, \pi)$ and $\omega_{e,\ell} \in [-\pi, \pi)$ are the beamspace angles (in radians) for the azimuth and elevation dimensions of the $\ell$-th path, respectively. $\mathbf{a}_N(\omega) \in \mathbb{C}^{N \times 1}$ is the array response Vandermonde vector for a given spatial frequency $\omega$.

25% masked · 1 blank · record #9721 · fill_blank_25

What is the missing array response vector for the complementary spatial dimension?

$\mathbf{H} = \sum_{\ell=1}^{L} \beta_\ell \mathbf{a}_N(\omega_{e,\ell}) \boxed{\,\color{#b3401c}{?}\,}$
Reveal answer

Ground truth: $\mathbf{a}_N^\mathrm{T}(\omega_{a,\ell})$

$\mathbf{H} = \sum_{\ell=1}^{L} \beta_\ell \mathbf{a}_N(\omega_{e,\ell}) \boxed{\color{#00706c}{\mathbf{a}_N^\mathrm{T}(\omega_{a,\ell})}}$
50% masked · 2 blanks · record #9722 · fill_blank_50

Fill in the complex path gain and the transposed array response vector.

$\mathbf{H} = \sum_{\ell=1}^{L} \boxed{\,\color{#b3401c}{?_{1}}\,} \mathbf{a}_N(\omega_{e,\ell}) \boxed{\,\color{#b3401c}{?_{2}}\,}$
Reveal answer

Ground truth: $?_{1} = \beta_\ell$, $?_{2} = \mathbf{a}_N^\mathrm{T}(\omega_{a,\ell})$

$\mathbf{H} = \sum_{\ell=1}^{L} \boxed{\color{#00706c}{\beta_\ell}} \mathbf{a}_N(\omega_{e,\ell}) \boxed{\color{#00706c}{\mathbf{a}_N^\mathrm{T}(\omega_{a,\ell})}}$
75% masked · 4 blanks · record #9723 · fill_blank_75

Complete the summation operator, the path gain, and the two angle variables.

$\mathbf{H} = \boxed{\,\color{#b3401c}{?_{1}}\,}_{\ell=1}^{L} \boxed{\,\color{#b3401c}{?_{2}}\,} \mathbf{a}_N(\boxed{\,\color{#b3401c}{?_{3}}\,}) \mathbf{a}_N^\mathrm{T}(\boxed{\,\color{#b3401c}{?_{4}}\,})$
Reveal answer

Ground truth: $?_{1} = \sum$, $?_{2} = \beta_\ell$, $?_{3} = \omega_{e,\ell}$, $?_{4} = \omega_{a,\ell}$

$\mathbf{H} = \boxed{\color{#00706c}{\sum}}_{\ell=1}^{L} \boxed{\color{#00706c}{\beta_\ell}} \mathbf{a}_N(\boxed{\color{#00706c}{\omega_{e,\ell}}}) \mathbf{a}_N^\mathrm{T}(\boxed{\color{#00706c}{\omega_{a,\ell}}})$
100% masked (FEC): the whole right-hand side · record #9724 · fill_blank_100

Write the complete equation for the physical channel model.

$\mathbf{H} = \boxed{\,\color{#b3401c}{?}\,}$
Reveal answer

Ground truth: $\sum_{\ell=1}^{L} \beta_\ell \mathbf{a}_N(\omega_{e,\ell}) \mathbf{a}_N^\mathrm{T}(\omega_{a,\ell})$

$\mathbf{H} = \boxed{\color{#00706c}{\sum_{\ell=1}^{L} \beta_\ell \mathbf{a}_N(\omega_{e,\ell}) \mathbf{a}_N^\mathrm{T}(\omega_{a,\ell})}}$
Derived from arXiv:2308.13268v1 records #9721–#9724 type fill_blank_25/50/75/100 train split CC BY 4.0 RedPajama-arXiv 13-gram hits: 0 · in S0 (all four)

Examples are drawn only from problems_cc_by.jsonl (records from CC BY 4.0 or CC0 source papers, released under CC BY 4.0) and shown verbatim with their released fields. Credit for the underlying mathematics belongs to the authors of the linked arXiv papers. Audit chips show each record's released per-problem metadata for the one fixed channel the audit covers (13-gram prompt-surface overlap against RedPajama-arXiv).

Chapter 3 · The reverse-probe audit

Index the small side, stream the big side

Auditing overlap at training-corpus scale means searching billions of corpus n-grams while keeping evidence per problem. The reverse probe flips the usual direction, so memory scales with the benchmark rather than the corpus. Scroll through six steps; the illustration follows along in four stages.

Reverse-probe audit, illustrated in four stages Stage 1: each benchmark prompt surface (background, question and masked equation, but not the answer) is normalised, cut into 13-token windows and hashed into a probe index that maps hashes to problem ids. Stage 2: the RedPajama-arXiv corpus is streamed through parallel workers while a 13-token window slides over its text. Stage 3: every corpus 13-gram is looked up in the index and each hit adds one to that problem's counter. Stage 4: problems with zero hits are labelled S0; problems whose hits come from at most two distinct documents are labelled S1. Rows and hashes are illustrative. 1INDEX THE BENCHMARK PROMPTS BACKGROUND QUESTION γ = MASK p17 answer + derivation: not indexed normalised tokens xxhash64 PROBE INDEX a3f9… → p17, p9020c71… → p4e52b… → p17 ~511K entries · ~50 MB 13-gram window 2STREAM THE CORPUS 13-GRAMS RedPajama- arXiv 1,558,306 records 16 parallel workers → 12.9 B 13-grams streamed · ~7 min · <6 GB peak 3COUNT PER-PROBLEM HITS each corpus 13-gram in the index? no: move on +1 PROBLEMHITS p4p17p902 021 matched n-grams kept (capped) for spot checks 4EMIT S₀ / S₁ LABELS p4 0 hits · 0 docs S₀ ✓ in S₁ ✓ in p17 2 hits · 1 doc S₀ ✗ out S₁ ✓ in p902 1 hit · 1 doc S₀ ✗ out S₁ ✓ in S₀: zero detected hits · S₁: hits from ≤ 2 distinct documents
Six steps, four stagesDiagram rows and hashes are illustrative.
  1. 1Stage 1 · index

    Take only what the model sees

    For each problem, the audit uses the released prompt surface: background, question text, and masked equation. It is normalised by lower-casing, stripping LaTeX command tokens, and collapsing whitespace.

    The unmasked correct_answer and source-paper derivation paragraphs are not part of the v1.0 audit input. So S0 is a prompt-surface lexical-overlap label, not a target-equation exposure certificate.

  2. 2Stage 1 · index

    Index the small side

    Every 13-gram of every prompt is hashed with xxhash64 into a probe index, a dictionary from gram hash to the problem IDs that contain it.

    ~511K unique entries~50 MB resident4,027 problems
  3. 3Stage 2 · stream

    Stream the big side

    The reference corpus is streamed file by file in parallel instead of being indexed. The release-defining target is RedPajama-arXiv: a 100-shard local snapshot of 1,558,306 records whose latest recorded timestamp is 7 March 2023.

    12.9 B 13-grams streamedmemory O(problems), not O(corpus)<6 GB peak · 16 workers · ~7 min
  4. 4Stage 3 · count

    Count hits, keep the evidence

    Each corpus 13-gram is looked up in the probe index. A hit increments that problem's counter and (capped at a small constant) records the matched n-gram for spot inspection.

    Dolma v1.7 arxiv and ProofPile-2 arxiv turned out to repackage the same RedPajama dump, so they add no independent in-domain evidence. OpenWebMath serves as an off-domain corroboration run: 99.87% of RedPajama-S0 problems also have zero detected OpenWebMath hits.

  5. 5Stage 4 · label

    Emit auditable labels

    The strict subset S0 keeps problems with zero detected RedPajama-arXiv matches; the lenient S1 allows up to two distinct source documents. Both definitions and their fallback rules were fixed before the audit ran; the fallbacks were not triggered.

    |S₀| = 3,853 (95.7%)|S₁| = 3,985 (99.0%)

    Hit counts, flags and memberships ship with the release, so users can tighten thresholds, audit new slices, or rebuild the filtered evaluation without regenerating the benchmark.

  6. 6Threshold

    Why n = 13, and why it matters

    Papers dated April 2023 or later cannot be flagged through their own text, so hits on them act as a collision-rate proxy. On the test split that proxy falls from 64.5% at n = 8 to 3.1% at n = 13. That rules out n = 8, but does not make 13 a validated optimum. The threshold was fixed before any full-vs-S0 result was examined.

    Zero-hit share of the 800-item test split by n:

    n = 833.0%
    n = 1075.0%
    n = 1293.1%
    n = 1396.2%
    n = 1598.8%

    S0 is therefore a fixed-threshold label, not a threshold-invariant property of a problem. The release exposes S0 membership for every tested n.

Chapter 4 · Reading the audit correctly

What the audit can say, and what it cannot

The audit gives a reproducible bound for exactly one channel. Everything outside that channel remains open, and we say so explicitly.

An auditable protocol for one fixed channel. Not a general cleanliness certificate.

In scope: one lit channel

  • One channel: exact 13-gram prompt-surface lexical overlap against RedPajama-arXiv at a fixed threshold.
  • Per-problem evidence: hit counts, S0/S1 flags, and S0 membership for each tested n, all recomputable.
  • A bounded stability summary: how much the few detected-overlap items could move the released scores (see the bound below).
  • Corroboration: an off-domain OpenWebMath run and an algebraic-stack pipeline check.

Outside this channel (not tested)

  • Paraphrase or near-duplicate exposure.
  • Target-answer exposure (answers and derivations are not indexed).
  • Alternate arXiv versions and closed-corpus mirrors.
  • Post-training exposure and reward-verifier co-adaptation.
  • Source-paper-disjoint generalisation: the v1.0 split is by problem (766 of 800 test items share a source paper with training).
  • A contamination-effect test. Full-vs-S0 is structurally underpowered for that.

Only 30 of 800 test items carry detected overlap

Waffle chart of the 800 test items: 770 with zero detected hits and 30 with detected overlap
  • 770 zero-hit items (S0, 96.2%)
  • 30 with detected overlap (buckets of 11 / 9 / 4 / 6)

One square per test item; the flagged items are grouped at the end for counting, not by position.

The bound

0.31–0.51 pp

If all 30 flagged items were counted as answered correctly (all-flagged-correct counterfactual), this is their largest possible positive score inflation for the five frontier rows.

So full-versus-S0 is a bounded, structurally underpowered stability summary, not a contamination-effect test or cleanliness claim.

The threshold is a design choice

At n = 8 only 264/800 test items stay zero-hit and the largest frontier point delta grows to +2.49 pp; for n ≥ 10 every frontier point delta stays within 0.96 pp. None of the ten frontier pairwise orderings changes sign across the full split and all five threshold views.

Frontier rows are a cluster, not a ranking

The five 16k frontier rows fall between 86.5% and 91.3%. The matched-budget GPT-5.2-chat vs GPT-5-mini delta is −0.50 pp [−2.88, +1.88]. Treat them as a high-accuracy calibration cluster, not a resolved rank order.

Judges and builders are disclosed

Cross-judge variation of the LLM-judge fallback is 4.6–7.3% absolute (per-model worst case), so smaller gaps are descriptive. DeepSeek-R1 (extraction) and GPT-4o (filtering, rubric scoring) helped build the benchmark; their rows are marked † and are descriptive only.

Chapter 5 · Results

Scores barely move when detected-overlap items are removed

Filtering to S0 changes accuracy by less than 1 pp for every evaluated model, across a 19-row calibration suite. Given the bound in Chapter 4, read this as stability under one audit channel, not as a leaderboard.

Frontier calibration at a 16k answer budget

Full = 800-item test split; S0 = 770-item strict zero-hit subset. Δ = Full − S0 with paired-bootstrap 95% CIs (10,000 resamples by problem id). Every row moves by at most 0.44 pp.

These five rows form one high-accuracy calibration cluster (86.5–91.3%), not a resolved rank order. Note the zoomed axis (84–94%).
  • Full test split (800)
  • Strict zero-hit subset S0 (770)
  • Δ = Full − S0, 95% paired-bootstrap CI
  • high-accuracy calibration cluster (86.5–91.3%)
  • GPT-5-miniFull 91.25 → S0 91.69
    Δ −0.44 pp, 95% CI [−1.05, +0.08]
  • GPT-5.2-chatFull 90.75 → S0 90.91
    Δ −0.16 pp, 95% CI [−0.67, +0.26]
  • GPT-4.1-miniFull 89.62 → S0 89.35
    Δ +0.27 pp, 95% CI [−0.03, +0.52]
  • Gemini-2.5-ProFull 86.88 → S0 87.01
    Δ −0.14 pp, 95% CI [−0.70, +0.33]
  • Gemini-2.5-FlashFull 86.50 → S0 86.49
    Δ +0.01 pp, 95% CI [−0.49, +0.43]

Locked 2k calibration rows (D1–D14)

Same full-vs-S0 protocol at a locked 2k budget. Within a row, full and S0 use the same budget; the 16k and 2k views are not a cross-budget leaderboard. † construction-involved, descriptive calibration only (excluded from comparative claims).

  • Full test split (800)
  • Strict zero-hit subset S0 (770)
  • Δ = Full − S0, 95% paired-bootstrap CI
General-purpose calibration models
  • D1Grok-4-FastFull 88.00 → S0 88.31
    Δ −0.31 pp, 95% CI [−0.90, +0.20]
  • D2DeepSeek-V3.1 (671B)Full 84.88 → S0 84.42
    Δ +0.46 pp, 95% CI [+0.14, +0.75]
  • D3Claude-4.0-SonnetFull 83.00 → S0 82.73
    Δ +0.27 pp, 95% CI [−0.20, +0.67]
  • D4o4-miniFull 81.88 → S0 82.08
    Δ −0.20 pp, 95% CI [−0.81, +0.34]
  • D5DeepSeek-R1 (671B)† (construction-involved, descriptive only)Full 79.12 → S0 79.09
    Δ +0.03 pp, 95% CI [−0.55, +0.55]
  • D6GPT-4o† (construction-involved, descriptive only)Full 77.00 → S0 76.88
    Δ +0.12 pp, 95% CI [−0.48, +0.63]
  • D7Qwen2.5-72B-Instruct (72B)Full 73.62 → S0 73.77
    Δ −0.14 pp, 95% CI [−0.80, +0.46]
  • D8Llama-3.3-70B-Instruct (70B)Full 61.50 → S0 61.56
    Δ −0.06 pp, 95% CI [−0.74, +0.60]
  • D9GPT-5-nanoFull 52.50 → S0 52.86
    Δ −0.36 pp, 95% CI [−1.07, +0.33]
Open 7B-scale and WirelessMathLM reference rows
  • D10WirelessMathLM-7B + GRPO (7B)oursFull 47.88 → S0 48.70
    Δ −0.83 pp, 95% CI [−1.51, −0.17]
  • D11DeepSeekMath-7B-RL (7B)Full 41.00 → S0 40.91
    Δ +0.09 pp, 95% CI [−0.58, +0.78]
  • D12Qwen2.5-Math-7B-Instruct (7B)Full 40.00 → S0 40.13
    Δ −0.13 pp, 95% CI [−0.77, +0.55]
  • D13WirelessMathLM-3B + GRPO (3B)oursFull 25.50 → S0 26.36
    Δ −0.86 pp, 95% CI [−1.28, −0.46]
  • D14WirelessMathLM-0.5B + GRPO (0.5B)oursFull 2.12 → S0 2.21
    Δ −0.08 pp, 95% CI [−0.14, −0.04]

Construction-involved. DeepSeek-R1 extracted the structured records and GPT-4o filtered papers and scored the QA rubric, so both helped build the benchmark. Their rows are retained only as descriptive calibration and are excluded from comparative claims.

The only systematic nonzero deltas are the WirelessMathLM rows (−0.83, −0.86, −0.08 pp; CIs exclude zero), in the direction opposite to what a simple prompt-surface lexical-overlap shortcut would predict. This is kept as a released sensitivity finding, not a causal claim.

WirelessMathLM rows are release-split learnability checks under the published verifier stack, not source-paper-disjoint or verifier-independent evidence.

Full 800 vs the 310-item public test subset

490 of the 800 test items are withheld under their source-paper licenses, so full-split numbers cannot be reproduced externally. Every row is re-scored on the 310-item public subset from the same archived per-item verdicts (no new model calls). CI is a 95% bootstrap over public items (10,000 resamples, seed 0); Δ = Public − Full.

  • Full test split (800)
  • Public test subset (310) with 95% CI
  • high-accuracy calibration cluster (86.5–91.3%)
Frontier rows (16k budget)
  • GPT-5-miniFull 91.25 → Public 310 90.32
    95% CI [86.77, 93.55] · Δ −0.93
  • GPT-5.2-chatFull 90.75 → Public 310 90.00
    95% CI [86.45, 93.23] · Δ −0.75
  • GPT-4.1-miniFull 89.62 → Public 310 89.35
    95% CI [85.81, 92.58] · Δ −0.27
  • Gemini-2.5-ProFull 86.88 → Public 310 85.48
    95% CI [81.61, 89.35] · Δ −1.39
  • Gemini-2.5-FlashFull 86.50 → Public 310 84.19
    95% CI [80.00, 88.06] · Δ −2.31
Locked-2k rows D1–D14
  • D1Grok-4-FastFull 88.00 → Public 310 84.84
    95% CI [80.65, 88.71] · Δ −3.16
  • D2DeepSeek-V3.1Full 84.88 → Public 310 83.87
    95% CI [79.68, 87.74] · Δ −1.00
  • D3Claude-4.0-SonnetFull 83.00 → Public 310 82.26
    95% CI [77.74, 86.45] · Δ −0.74
  • D4o4-miniFull 81.88 → Public 310 79.35
    95% CI [74.84, 83.87] · Δ −2.52
  • D5DeepSeek-R1† (construction-involved, descriptive only)Full 79.12 → Public 310 77.74
    95% CI [72.90, 82.26] · Δ −1.38
  • D6GPT-4o† (construction-involved, descriptive only)Full 77.00 → Public 310 75.81
    95% CI [70.97, 80.32] · Δ −1.19
  • D7Qwen2.5-72B-InstructFull 73.62 → Public 310 73.87
    95% CI [69.03, 78.71] · Δ +0.25
  • D8Llama-3.3-70B-InstructFull 61.50 → Public 310 59.35
    95% CI [53.87, 64.84] · Δ −2.15
  • D9GPT-5-nanoFull 52.50 → Public 310 51.29
    95% CI [45.81, 56.77] · Δ −1.21
  • D10WirelessMathLM-7BoursFull 47.88 → Public 310 46.45
    95% CI [40.65, 51.94] · Δ −1.42
  • D11DeepSeekMath-7B-RLFull 41.00 → Public 310 38.71
    95% CI [33.23, 44.19] · Δ −2.29
  • D12Qwen2.5-Math-7B-InstructFull 40.00 → Public 310 40.65
    95% CI [35.16, 46.13] · Δ +0.65
  • D13WirelessMathLM-3BoursFull 25.50 → Public 310 25.16
    95% CI [20.32, 30.00] · Δ −0.34
  • D14WirelessMathLM-0.5BoursFull 2.12 → Public 310 2.26
    95% CI [0.65, 4.19] · Δ +0.13

Construction-involved. DeepSeek-R1 extracted the structured records and GPT-4o filtered papers and scored the QA rubric, so both helped build the benchmark. Their rows are retained only as descriptive calibration and are excluded from comparative claims.

Among the five frontier rows no pairwise ordering changes; among the 14 locked-2k rows, one of 91 pairs changes order (DeepSeekMath-7B-RL vs Qwen2.5-Math-7B-Instruct, 1.00 pp apart on the full split). The public subset is slightly harder for most rows, and it is the recommended view for externally comparable results.

Learnability check

A size-matched Math12K control

Same Qwen2.5-7B base, 3,227 training examples, 40 epochs / 240 steps, locked 800-item evaluation; one run per condition.

Math12K GRPO control22.75%
WirelessMathLM-7B47.88%

Paired Δ −25.13 pp, 95% item-bootstrap CI [−29.00, −21.13]. This supports release-split domain-aligned learnability, not intrinsic difficulty, causal isolation of wireless constraints, seed robustness, or paper-disjoint generalisation. Lower-level batch, prompt-template and reward settings are not identical.

Learnability check

GRPO from base checkpoints

Training-time checks under greedy model selection (separate from the locked T = 0.6 + judge scores).

Qwen2.5-3B base12.37%
Qwen2.5-3B + GRPO25.12%
Qwen2.5-7B base21.88%
Qwen2.5-7B + GRPO39.50%

The 0.5B checkpoint moves only +1.49 points and stays near the floor; a single Qwen3-4B run improves by +4.24. These rows show the training split is learnable with a verifier reward under the published stack; the shared GPT-4.1-mini verifier family keeps verifier-independent reasoning outside the claim.

Chapter 6 · Open release

Take it, re-run it, tighten it

Redistribution follows each source paper's license (harvested from arXiv OAI-PMH). Records whose license does not clearly permit derived text are withheld and listed by identifier only. Identifier-level audit metadata for all 4,027 problems is released under CC BY 4.0; code is MIT-licensed.

What is public

All 4,027 records1,542 distributed
800-item test split310 distributed
  • CC BY 4.0 (CC BY / CC0 sources): 1,334 · test 273
  • CC BY-SA 4.0: 43 · test 9
  • CC BY-NC-SA 4.0: 165 · test 28
  • Withheld (ID only): 2,485 · test 490
Report results on the 310-item public test subset for externally comparable numbers; it is the union of the three test splits on the Hub (273 + 9 + 28), and its problem IDs also ship as public_test_ids.json. Reference scores for all 19 rows are in the Public 310 tab.

Load it

Planned Hugging Face layout · three license configs
ConfigLicenseRecordsIn test
cc_by (default)CC BY 4.01,334273
cc_by_saCC BY-SA 4.0439
cc_by_nc_saCC BY-NC-SA 4.016528

Each config has train and test splits. The 2,485 withheld records are listed by ID only in excluded_records_manifest.csv.

# pip install datasets
from datasets import load_dataset, concatenate_datasets

REPO = "XINLI1997/WirelessMATHBench-XL"
test = load_dataset(REPO, "cc_by", split="test")

# 310-item public subset = the three test splits
public = concatenate_datasets([
    load_dataset(REPO, cfg, split="test")
    for cfg in ("cc_by", "cc_by_sa", "cc_by_nc_sa")
])

item = public[0]
print(item["type"], item["paper_id"])
print(item["prompt"])          # what the model sees
print(item["correct_answer"])  # verifier-facing truth
print(item["contamination"])   # per-problem audit metadata

In the release

  • problems_cc_by*.jsonldistributed records by license (1,334 / 43 / 165), mirrored as the three Hub configs
  • excluded_records_manifest.csv2,485 withheld records, identifiers only
  • public_test_ids.jsonthe 310-item public test subset
  • contamination_metadata.jsonlhit counts and S0/S1 flags for all 4,027
  • s0_membership_by_n.csvS0 membership for each tested n
  • paper_disjoint_test34-item source-paper-disjoint sensitivity view (+ seed-0 grouped 3,222 / 805 split)

Also: evaluation traces, paired-bootstrap scripts, GRPO training recipes, Croissant metadata, and a Datasheet for Datasets. File names follow the release data folder.

Chapter 7 · Citation

Cite WirelessMathBench-XL

If you use the benchmark, the audit metadata, or the reference checkpoints, please cite:

@inproceedings{
li2026wirelessmathbenchxl,
title={WirelessMathBench-XL: An Auditable Benchmark for Wireless Mathematical Reasoning},
author={Xin Li and Mengbing Liu and Yiyang Zhu and Wenhe Zhang and Li Wei and Jiancheng An and Chau Yuen},
booktitle={The Fortieth Annual Conference on Neural Information Processing Systems Evaluations and Datasets Track},
year={2026}
}