iChancellor’s AI
← All action models
Open model

Chancellor-VLM 40B

Multimodal: text + images, openly

40B params64K tokens + 8 imagesApache-2.0Launches November 11, 2026

Not released yet — every artifact below ships on launch day. Until then, the Actionfield preview simulates the experience.

About this model

Chancellor-VLM reads figures, lab plots, documents, and photographs alongside text. It was built for science: extracting structure from charts, reasoning about experimental imagery, and reading the tables in papers.

Its vision corpus is the part we're proudest of. Multimodal training data is usually scraped first and questioned later; ours carries the same full audit as our language corpora — provenance and consent documentation for every image source. It took a year longer to assemble. It was worth it.

Architecturally, the model couples a vision encoder to the Chancellor action backbone with a fusion layer trained to preserve spatial detail — which is why it can tell you not just that a curve rises, but where the knee is and what the axis labels say. Up to eight images share a 64K-token context, so it can compare figures across a whole paper.

When text and visual evidence conflict, Chancellor-VLM is trained to flag the disagreement and state which source it weighted. In evaluation, that single behavior eliminated most silent hallucinations on chart-reading tasks.

What it’s good at

  • Chart, plot, and table extraction from papers
  • Visual reasoning over experimental imagery
  • Document OCR with layout understanding
  • Flags text-image disagreement instead of guessing
  • Consent-documented, fully audited vision corpus

Where people use it

Extracting datasets from published figuresLab-notebook and instrument-readout analysisAccessible descriptions of scientific imageryDocument digitization with layout fidelity
Method

How it was trained — in the open

01

Consent-first corpus

Every image source carries provenance and consent documentation — audited and published, a first at this scale.

02

Vision-language fusion

A spatial-detail-preserving fusion layer trained on scientific figures, not just web photos.

03

Grounded instruction tuning

Chart-reading demonstrations paired with the underlying data tables, so answers ground in values, not vibes.

04

Disagreement training

Deliberately conflicting text-image pairs teach the model to surface contradictions rather than smooth over them.

Evidence

Benchmarks, with context

Numbers without baselines are marketing. Every score below ships with its comparison point and its caveat.

BenchmarkThis modelReferenceNote
SciChart-QA (chart reading)78.974.2 · best open VLM peergrounded in axis values
DocTable extraction (F1)91.388.0 · best open VLM peerlayout-aware OCR
LiveEval-V (contamination-resistant)58.453.9 · best open VLM peerimages postdate cutoff
Silent hallucination rate3.1%11.7% · open VLM averagelower is better
Open artifacts

Everything ships at launch

Launches November 11, 2026. Want a note the moment they’re live? Get notified.

Model weights

safetensors · 80 GBLaunches November 11, 2026

Vision data audit

provenance + consentLaunches November 11, 2026

Training code

Apache-2.0Launches November 11, 2026

VQA evaluation set

science-focusedLaunches November 11, 2026