H1 2026 safety report
What we found red-teaming Cairn-3 for six months, and the mitigations that made it into the release.
Subscribe to running log of model releases, product updates, and research notes from the team building Cairn.
What we found red-teaming Cairn-3 for six months, and the mitigations that made it into the release.
Repeated context is cached automatically across requests — up to 90% cheaper and 4x faster for long prompts.
Talk to Cairn with sub-second latency and interruptions that feel natural, not scripted.
The harness we use to measure reasoning, honesty, and tool use — free for anyone to run against any model.
Connect a repo and get inline review comments that understand your codebase, not just the diff.
Guarantee JSON that matches your schema on every response — no more retry loops around malformed output.
Queue non-urgent workloads and get results within 24 hours at 50% off standard rates.
New interpretability tooling lets you trace which inputs shaped a given output, down to the token.
Up to 1M tokens of context, with retrieval quality that holds steady all the way to the edge.
Our latest model calls tools mid-reasoning instead of waiting for a final answer — faster agents, fewer dropped steps.
Queue non-urgent workloads and get results within 24 hours at 50% off standard rates.
New interpretability tooling lets you trace which inputs shaped a given output, down to the token.
Up to 1M tokens of context, with retrieval quality that holds steady all the way to the edge.
Our latest model calls tools mid-reasoning instead of waiting for a final answer — faster agents, fewer dropped steps.
What we found red-teaming Cairn-3 for six months, and the mitigations that made it into the release.
Repeated context is cached automatically across requests — up to 90% cheaper and 4x faster for long prompts.
Talk to Cairn with sub-second latency and interruptions that feel natural, not scripted.
The harness we use to measure reasoning, honesty, and tool use — free for anyone to run against any model.