Chancellor-1 70B Instruct
Our flagship open Large Action Model
Not released yet — every artifact below ships on launch day. Until then, the Actionfield preview simulates the experience.
About this model
Chancellor-1 is a fully open 70-billion-parameter Large Action Model (LAM) built for scientific work. It is not a language model that happens to call tools — it is trained end-to-end to plan, decide, and execute multi-step actions, with text as one of its instruments rather than its whole world. And unlike closed frontier models, every part of it is public: the weights, the 4.1T-token training data recipe, the training code, and every intermediate checkpoint.
The design goal was never raw benchmark scores — it was accountable capability. Chancellor-1 was trained with a complete data audit, which means every ability it demonstrates can be traced back to the corpus slices that produced it. That audit is what powers ChancellorTrace in the Actionfield: for any response, you can inspect training documents with exact text matches.
The instruct variant went through three alignment stages, each documented publicly: supervised fine-tuning on openly licensed demonstrations, preference tuning against a released reward model, and a final calibration pass that teaches the model to express uncertainty rather than bluff. The result is a model that says "I'm not sure" measurably more often when it is, in fact, not sure.
Because all 47 intermediate checkpoints are published, Chancellor-1 doubles as a scientific instrument: researchers use the checkpoint series to study how reasoning ability emerges during training — work that is simply impossible with closed models.
What it’s good at
- →Graduate-level science questions and multi-step reasoning
- →Tool calling with open, auditable tool schemas
- →Long-document analysis up to 128K tokens
- →Calibrated uncertainty — it knows what it doesn't know
- →Full reproducibility — retrain it yourself from released artifacts
Where people use it
How it was trained — in the open
Corpus assembly & audit
4.1T tokens gathered under documented licenses, deduplicated, and indexed token-by-token so every capability can be traced to its sources.
Pretraining
Trained across 47 published checkpoints with full loss curves and data-order logs — the entire run is replayable.
Supervised fine-tuning
Instruction demonstrations from openly licensed and consented sources only; the SFT set itself is downloadable.
Preference tuning
Aligned against a released open reward model, so the values baked in are inspectable, not implied.
Calibration pass
A final stage rewards honest uncertainty, cutting confident-but-wrong answers by a third on held-out evals.
Benchmarks, with context
Numbers without baselines are marketing. Every score below ships with its comparison point and its caveat.
| Benchmark | This model | Reference | Note |
|---|---|---|---|
| GradSci-QA (graduate science) | 71.4 | 73.1 · closed frontier | within 2 points, fully auditable |
| LiveEval (contamination-resistant) | 64.8 | 58.2 · best open peer | items postdate training cutoff |
| MathBench | 68.2 | 66.9 · best open peer | chain-of-thought, no tools |
| ToolUse-Hard | 82.1 | 79.4 · closed frontier | open tool schemas |
Everything ships at launch
Launches September 9, 2026. Want a note the moment they’re live? Get notified.