Notes · Reading list · verifiable science data
Ten papers on scalable, verifiable science data and environments for the Apertus 2 STEM climb. Nine are from the last twelve months; only MegaScience (July 2025) is older. Grouped by where the verifiable signal comes from.
The four for the slide: one executable environment, one simulator, one mined corpus, and the held-out target. Each in plain words first, then the numbers.
Papers and textbooks become multiple-choice or short-answer questions. Cheap, broad, verifiable by string match; quality rests on the filtering.
Answers are computed, not annotated. Unlimited supply and controllable difficulty; the question is transfer to real problems.
Agentic science: the model writes and runs code against real data, a hidden verifier grades. Closer to the tools/SWE climb than to a reasoning climb; the transferable idea is verifier synthesis.
Generation infrastructure rather than a dataset.
Caveat for any GPQA claim above: HLE-Physics moved from 47 to 79 after expert regrading (arxiv 2609.13009), closed-ended physics evals are noisy.
How the frontier reports scope the reasoning, STEM or math expert. Only MAI-Thinking-1 publishes the recipe; the rest give a scope line and the evals. Read for the Apertus 2 STEM climb (no issue yet), where the plan is copied from MAI.
| Report | Expert | Scope | Data and reward | Tools | Evals |
|---|---|---|---|---|---|
| MAI-Thinking-1 §3.2 | STEM climb, 1 of 3 specialists | single-turn problem solving: math, physics, chemistry, biology, CS, engineering, plus competitive code | STEM Mix, 5M (q, a) pairs mined from textbooks, PDFs, forums, contests, vendors; 550K hard; SymPy or judge, test cases for code | no | AIME, HMMT, GPQA-D, LCB v6 |
| Nemotron 3 Ultra §3.3.2 | STEM / general-reasoning teacher, 1 of 11 | math, code, natural sciences, humanities, sociology | SFT 40B tokens (science 59%, math 24%, code 10%, general 7%), traces from DeepSeek-V4-Pro judged by gpt-oss-120b; then RL on non-STEM prompts, 128 per batch | "for these domains" | HLE, GPQA, MMLU-Pro, LCB v6, IMOAnswerBench, Apex Shortlist |
| Nemotron 3 Super | none, one joint RLVR policy | 21 environments: math with and without Python, formal proofs, competition code, STEM, IF, safety, long context, agentic | prompts the SFT model always solves are dropped; joint training, single-environment runs regress | Python tool for math | AIME, GPQA-D, LCB, plus agentic |
| Nemotron-Cascade 2 | math teacher = the SFT checkpoint | competition math and proofs | SFT: 2.6M math samples, 1.8M with tool calls, 816K proofs; MOPD from math, RLHF and multi-domain teachers | yes in SFT | AIME 25/26, HMMT, IMOAnswerBench, IMO-ProofBench |
| Kimi K3 §4.1.2 | none: a "general" expert | general experience, vision, reasoning, faithfulness, search, knowledge work | verifiable problems in agentic environments plus an agentic generative reward model; 3 domains × 3 effort levels | yes | knowledge and reasoning suites, HLE |
| DeepSeek-V4 §5.1 | a mathematics expert, 1 of 10+ | not described | SFT then GRPO on "domain-specific prompts and reward signals"; rule-based verifiers, generative RM for the rest; 3 effort modes per expert | Lean 4 agentic for formal math | HMMT, IMOAnswerBench, Apex, HLE, GPQA-D, Putnam |
| Qwen3 §4 | none: Reasoning RL stage of one model | math and code | 3,995 query-verifier pairs, unseen in cold start, learnable, hard; GRPO 170 steps | no | AIME'24 70.1 to 85.1 (235B) |
| AceReason-Nemotron 1.1 | none: math-only then code-only RL | math, code | prompt diversity beats responses per prompt; SFT gap 6.6 to 1.6 AIME after RL | no | AIME 24/25, LCB v5/v6 |
| GLM-5 | none: Reasoning RL stage, then agentic, general, cross-stage OPD | not described | not described | ||
| Apertus 1.5 | none: one RLVR stage | math and code | 349,300 prompts, async verl, then DPO and SDPO | no | |
| Apertus 2 plan | "STEM math and general reasoning", 1 of 7 climbs | undefined; AIME-level target, grade-school math as prerequisite; Lean and code under #1027 | TBD | open question | Artificial Analysis Index |