Built from an Olmo 3 run that got 21 extra days of reinforcement learning, Olmo 3.1 32B Think is for researchers and teams who must audit or retrain a reasoning model end to end. Pick it for transparency, not peak capability: newer open models report far higher GPQA Diamond scores.
Olmo 3.1 32B Think is a 32.2-billion-parameter dense reasoning model from Ai2, released with Apache 2.0 weights plus its training data, code and checkpoints. Ai2 reports 78.1 on AIME 2025 and 96.2 on MATH. Artificial Analysis scores it 7 on its own index, so it suits research and auditing more than frontier work.
Where it sits
- 57.1%% correctGPQA DiamondHigher is better#46 / 49peer median 88.9%per source, see benchmark scores
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: Ai2 · Family: Olmo 3
Context window: 65,536 tokens
Input modalities: text · Output: text
About Olmo 3.1 32B Think
Olmo 3.1 32B Think is the reasoning variant of Ai2's Olmo 3 family, announced on December 12, 2025 as an update to Olmo 3 32B Think from November 2025. Ai2 resumed the reinforcement learning run for 21 more days on 224 GPUs with extra passes over its Dolci-Think-RL data, and reports gains of 5 or more points on AIME, 4 or more on ZebraLogic and IFEval, and 20 or more on IFBench. It is a dense, decoder-only Transformer with 32,233,522,176 parameters, 64 layers, a 5,120-wide hidden state and 8 key-value heads shared across 40 attention heads. Three of every four layers use a 4,096-token sliding attention window and every fourth layer attends to the full context, which Ai2 extended to 65,536 tokens.
The model is fully open in a stricter sense than most open-weight releases. Ai2 publishes the weights under Apache 2.0, the Dolma 3 pretraining corpus (about 9.3 trillion tokens of web pages, olmOCR-processed science PDFs, code, math and encyclopedic text), a 5.9-trillion-token pretraining mix, 100 billion tokens of mid-training data, about 50 billion tokens for long-context extension, and the Dolci sets used for supervised fine-tuning, preference tuning and reinforcement learning. Pretraining used up to 1,024 H100 GPUs, and the data cutoff is December 2024. In the Ai2 Playground, the OlmoTrace tool can trace parts of a response back to matching training documents.
Ai2's model card, checked on October 3, 2026, lists 78.1 on AIME 2025, 80.6 on AIME 2024, 96.2 on MATH, 86.4 on MMLU, 83.3 on LiveCodeBench v3, 91.5 on HumanEvalPlus, 80.1 on ZebraLogic, 93.8 on IFEval and 68.1 on IFBench. In Ai2's own table it beats Qwen 3 32B on AIME 2025 (70.9) and IFBench (37.3) but trails it on GPQA (67.3), LiveCodeBench v3 (90.2) and ZebraLogic (88.3). IFBench is an instruction-following test that Ai2 created, so read that gap with care.
Independent figures are thin. Artificial Analysis scores the model 7 on its Intelligence Index v4.3.2 (checked October 3, 2026) and publishes no speed or cost data for it. Ai2's own GPQA number differs between its sources: 56.7 in the technical report and 57.5 on the model card, while a community evaluation record on the Hugging Face page lists 57.07 on GPQA Diamond, the figure stored here. Phi-4, a 14B model, reports a similar 56.1 on GPQA Diamond, while newer releases like Qwen3.8 27B report 89.2 and Gemma 4 31B reports 85.7 on the same test, each vendor-reported.
The weights are free to download. The BF16 checkpoint is 64.5 GB across 14 safetensors files, so the weights alone need an 80 GB accelerator or a multi-GPU split before any KV cache. Transformers 4.57.0 or later, vLLM and SGLang can serve it, and Hugging Face lists 25 community quantizations. Ai2 recommends temperature 0.6, top-p 0.95 and a 32,768-token output budget. Ai2 offered the model in its Playground at launch, and its December 2025 post said inference partners would host it, but on October 3, 2026 OpenRouter showed no live endpoints and Artificial Analysis listed one host, Parasail, with no verified price.
Ai2 reports a composite safety score of 83.6 for this model, against 69.0 for Qwen 3 32B in the same table. The technical report says the development safety suite includes HarmBench, DoAnythingNow, XSTest, WildGuardTest, WildJailbreak and TrustLLM-JailbreakTrigger. The model card warns that, without safety filtering, the model can be prompted into harmful content, and the license note says it is intended for research and educational use under Ai2's Responsible Use Guidelines. No third-party red-team partner is named.
Choose it to study or reproduce reasoning-focused reinforcement learning, to audit what a model was trained on, or to self-host an Apache 2.0 reasoning model whose lineage you can document. Skip it for long documents (65,536 tokens against 262,144 on both Qwen3.8 27B and Gemma 4 31B, or 1,000,000 on NVIDIA's Nemotron 3.5 Lightning), for non-English work (the card lists English) and for agent products that need current frontier accuracy. Ai2's Olmo 3.1 32B Instruct targets chat and tool use instead. We found no newer Olmo language model as of October 3, 2026, and Ai2's Olmo co-lead Hanna Hajishirzi was reported in March 2026 to be joining Microsoft, so treat future updates as uncertain. More choices sit in the open-source models directory, and the guides Which LLM Should You Use? and Best AI Models You Can Run Locally cover picking and self-hosting.
Pricing
The weights are free under a permissive license, so the cost is the hardware to run a 32B dense model. Ai2 has not published a first-party API price. Artificial Analysis lists one host, Parasail, with a $0.00 blended price that we treat as unconfirmed. For a priced open alternative, [Qwen3.8 27B](/hub/models/qwen3.8-27b) is sold by Alibaba Cloud at $0.50 for input and $3.00 for output, per million tokens.
| Tier | Rate |
|---|---|
| Self-hosted weights | Free download under Apache 2.0; you pay for hardware |
Key Features
- Fully open model flow: Weights, Dolma 3 pretraining data, Dolci post-training data, training code and intermediate checkpoints are all downloadable.
- Hybrid attention: Most layers look only at nearby text, while every fourth layer sees the entire window, which keeps long inputs affordable.
- Extended reinforcement learning: Ai2 kept training with reinforcement learning after the Olmo 3 launch, which it credits for a gain of more than 20 points on IFBench.
- OlmoTrace: In the Ai2 Playground, parts of an answer can be traced to matching documents in the training data.
- Permissively licensed weights: A single checkpoint that serves with Transformers, vLLM or SGLang, with no usage-threshold clauses.
Pros
- Best fit when the goal is research or audit: every stage is reproducible, which most open-weight releases of this size do not offer.
- Permissive license with no usage-threshold clauses, and a dense design that is easier to serve than a mixture of experts.
- Honest vendor documentation: the technical report discloses decontamination work and per-stage results.
Cons
- Low on independent measures: Artificial Analysis scores it 7 on its Intelligence Index v4.3.2.
- Short window next to newer open models, and text-only input.
- Hosted access is uncertain, so most teams must run their own GPUs.
Benchmarks
- MATH: 96.2% vendor-reported · 03 Oct 2026 — Competition maths problems, % solved.
- MMLU: 86.4% vendor-reported · 03 Oct 2026 — General-knowledge exam across 57 subjects, % correct.
- IFBench: 68.1% vendor-reported · 03 Oct 2026 — How precisely the model follows detailed instructions, % passing.
- AIME 2025: 78.1% vendor-reported · 03 Oct 2026 — Competition-level maths problems from the 2025 exam, % solved.
- GPQA Diamond: 57.1% independent · 03 Oct 2026 — PhD-level science questions that are hard to search for, % correct.
- AA Intelligence Index: 7 cited: Artificial Analysis · 03 Oct 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.
Frequently Asked Questions
What does Olmo 3.1 32B Think cost to run?
Downloading the weights is free, so the real cost is hardware: the files total 64.5 GB. Ai2 has not published a first-party API price, and hosted access is patchy: OpenRouter listed no live endpoints on October 3, 2026, and Artificial Analysis showed a single host with a $0.00 blended price that we could not confirm. Budget for your own GPUs or a rented node rather than a per-token bill.
How does Olmo 3.1 32B Think compare with Qwen 3 32B and newer open models?
In Ai2's own table it leads Qwen 3 32B by 7.2 points on AIME and 30.8 points on IFBench, a test Ai2 wrote, but trails by about 10 points on GPQA and 6.9 on LiveCodeBench v3. Newer open models are well ahead on science questions: Qwen3.8 27B reports 89.2 on GPQA Diamond and Gemma 4 31B reports 85.7, both vendor-reported. Pick Olmo when training transparency matters more than the highest score.
Is Olmo 3.1 32B Think open source?
The weights carry a permissive Apache license, and Ai2 also releases the training data, code and intermediate checkpoints, which is what it means by fully open. The datasets are downloadable without license restrictions, according to Ai2's launch post. The model card also asks that use follow Ai2's Responsible Use Guidelines, which frame the release around research and teaching and are worth reading before a commercial launch.
Does Olmo 3.1 32B Think train on your data?
Self-hosted copies send nothing back to Ai2, so your prompts stay on your own machines. We found no model-specific retention policy for Ai2's hosted Playground, and the site privacy notice covers website data and predates this model's release. If your prompts are sensitive, run the weights locally.
Who should use Olmo 3.1 32B Think, and who should pass?
It fits researchers studying reasoning-focused reinforcement learning, auditors who need to see what a model learned from, and teams that want a permissively licensed reasoning model they can fully document. Pass if you need documents longer than its window, languages other than English, or current frontier accuracy: a newer open model such as Qwen3.8 27B or Gemma 4 31B suits those better. For chat and tool use, Ai2 points to its separate Instruct model.
Top Alternatives
- Qwen3.8 27B: Pick Olmo 3.1 32B Think for published training data; pick Qwen3.8 27B for a much longer window and a far higher vendor-reported GPQA Diamond score.
- Gemma 4 31B: Pick Olmo to audit or retrain the pipeline; pick Gemma 4 31B for image input and a longer window.
- Phi-4: Pick Olmo for a longer window and open training data; pick Phi-4 for a smaller 14B footprint under an MIT license.