Ornith-1.5: The Open Model That Beats Claude
Most model releases follow the same script. Bigger model, more data, a few points on the leaderboard, a blog post full of adjectives.
Ornith-1.5 is doing something else, and the something else is worth understanding even if you never download the weights.
During training, this model writes its own practice problems. It builds the grading scripts for those problems. It attempts them, and the results feed back into all three of those steps at once. There is no fixed pile of human-written tasks sitting underneath it. The curriculum generates itself and gets harder as the model gets better.
That’s the pitch. Here’s what’s actually in the release, what the numbers show, and where the framing runs ahead of the data.
What shipped
Three sizes, all MIT licensed:
| Model | Type | Notes |
|---|---|---|
| Ornith-1.5-397B | MoE | The flagship |
| Ornith-1.5-35B-A3B | MoE | Roughly 3B parameters active per token |
| Ornith-1.5-9B | Dense | Single-GPU, with a quantized Mobile variant |
All three extend Ornith-1.0, which the team built on top of Qwen3.5 and Gemma 4 with additional continued pretraining, mid-training, and post-training.
The 9B has a quantized sibling called Ornith-1.5-9B-Mobile that the blog says targets iPhone and Android. The blog doesn’t give a file size for it, so I’m not going to invent one.
How the self-improvement loop works
Training a coding model normally means humans write thousands of problems and the test scripts that verify solutions. The model grinds against that fixed set. It works, but you eventually run out of problems, and the ones you wrote may not target what the model is actually bad at.
Ornith-1.5 removes the fixed set. Each training cycle runs three stages.
The model proposes a task. It gets an environment or codebase, high-level guidance about the kind of task to produce, and its own history of what it has already solved. From that it writes a harder problem, aimed at a gap in its own ability.
The model builds a scaffold. In Ornith’s terms, the scaffold is the instructions, the tools, the decomposition strategy, and the orchestration used to approach the problem. Not a prompt template. The whole apparatus around the task.
The model attempts a solution. Conditioned on both the task and the scaffold, the policy produces a rollout.
Then reward from that rollout propagates back through all three stages. The model isn’t only learning to solve better. It’s learning to write more useful tasks and build more effective scaffolds, in the same loop.
Run that repeatedly and stronger policies produce harder tasks, harder tasks produce better training signal, and the scaffolds keep evolving toward whatever actually elicits the model’s capability.
The part that keeps it honest
If a model invents its own homework, the obvious failure mode is that it invents easy homework and farms a great score.
Ornith scores every generated task on three signals and multiplies them together, so a task has to satisfy all three or the reward collapses.
Validity. Is the task coherent and solvable? Does the scaffold execute correctly and evaluate candidates reliably? The checks include whether high-confidence solutions pass and clearly incorrect ones fail. This is a hard gate. If validity comes back zero, the entire task reward is zero, which stops malformed tasks from getting paid just for looking difficult.
Frontier difficulty. They sample multiple rollouts per task and measure the empirical success rate, then reward tasks whose rate sits near a target. That target is set at 0.2. Roughly one attempt in five succeeds, which is hard enough to be worth learning from and easy enough to produce usable successful trajectories. The elegant part: as the model starts reliably solving a task, that task’s reward drops on its own, pushing the generator toward harder ones.
Novelty. Each task is compared against a buffer of previously generated and trained-on tasks. Too similar, and novelty drops. Without this term, frontier difficulty alone would encourage a thousand slight variations of the same problem.
There’s a parallel reward for the harness itself, scored on task alignment, reward fidelity, and hack resistance. All three stages are optimized with GRPO.
The benchmarks
397B
The flagship is the one trading blows at the top.
| Benchmark | Ornith-1.5-397B | Claude Opus 4.8 |
|---|---|---|
| Terminal-Bench 2.1 (Terminus-2) | 86.1 | 85.0 |
| SWE-bench Verified | 86 | 85.8 |
| DeepSWE | 56 | 59 |
It takes Terminal-Bench and SWE-bench Verified by narrow margins and loses DeepSWE. “Matches” is the accurate word. “Beats” isn’t.
35B
The 35B is arguably the most interesting one in the family. It activates only about 3B parameters per token and still posts 68.5 on Terminal-Bench 2.1 against Gemma-4-31B’s 43.4 and Muse-Glimmer-30B’s 51.7. On SWE-bench Verified it hits 79.0 against 52.0 and 76.0 respectively.
It also beats Qwen3.6-35B-A3B across the coding and agentic benchmarks, which is a like-for-like comparison at the same scale.
9B
This is the one most people can actually run.
| Benchmark | Ornith-1.5-9B | Ornith-1.0-9B | Gemma-4-31B | Qwen3.6-35B-A3B |
|---|---|---|---|---|
| Terminal-Bench 2.1 (Terminus-2) | 46.2 | 43.1 | 42.1 | 52.5 |
| Terminal-Bench 2.1 (Claude Code) | 47 | 40.6 | – | 49.2 |
| SWE-bench Verified | 70.6 | 69.4 | 52 | 73.4 |
| SWE-bench Pro | 47.5 | 42.9 | 35.7 | 49.5 |
| GPQA Diamond | 86.4 | 82.5 | 84.3 | 86 |
| BrowseComp | 56.4 | 44.8 | – | 62 |
| ClawEval | 66.5 | 63.1 | 48.5 | 68.7 |
Against Gemma-4-31B the 9B wins clearly, and SWE-bench Verified at 70.6 against 52 is a wide gap in favor of a model less than a third the size.
The BrowseComp jump is the one I’d flag as most notable. 44.8 to 56.4 in a single version is a big move.
Where the framing overshoots
The blog says the 9B substantially outperforms larger models including Gemma-4-31B and Qwen3.6-35B.
The first half holds. The second doesn’t, and their own table shows it. Qwen3.6-35B-A3B beats the 9B on Terminal-Bench (52.5 to 46.2), SWE-bench Verified (73.4 to 70.6), SWE-bench Pro (49.5 to 47.5), BrowseComp (62 to 56.4), and ClawEval (68.7 to 66.5).
The 9B does win on NL2Repo (32.4 to 29.4), SWE Atlas QnA (20.6 to 15.5), and edges GPQA Diamond (86.4 to 86). So it’s a real fight against a model roughly four times its size, which is genuinely impressive. It just isn’t a sweep, and publishing the table that contradicts your own summary sentence is an odd choice.
Worth crediting the methodology, though. Results are averaged over five independent runs. For SWE-bench evaluations they stripped git history out of the repository images so the model couldn’t look up the commit that fixed the issue, and disabled network access so it couldn’t retrieve the answer. Those are the safeguards you’d want to see.
These are still self-reported numbers from the team’s own evaluation setup. Normal for a launch, but worth holding lightly until others reproduce them.
Running the 9B
The model card is specific about requirements. Ornith-1.5-9B is roughly 19 GB in bf16 and serves on a single 80GB GPU.
You’ll need recent runtimes: Transformers 5.8.1+, vLLM 0.19.1+, or SGLang 0.5.9+.
vLLM:
bash
vllm serve ornith-ai/Ornith-1.5-9B \
--served-model-name Ornith-1.5-9B \
--host 0.0.0.0 --port 8000 \
--max-model-len 262144 \
--gpu-memory-utilization 0.90 \
--enable-prefix-caching \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \
--trust-remote-code
SGLang:
bash
python -m sglang.launch_server \
--model-path ornith-ai/Ornith-1.5-9B \
--served-model-name Ornith-1.5-9B \
--host 0.0.0.0 --port 8000 \
--context-length 262144 \
--mem-fraction-static 0.85 \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3
Ollama:
bash
ollama run ornith-1.5:9b
Two things to know before you start sending requests.
It’s a reasoning model. The assistant turn opens with a thinking block before the final answer. The serving recipes above enable a reasoning parser so that trace comes back in a separate reasoning_content field, which means you can inspect it during development and hide it in production without string-parsing anything.
Use their sampling parameters, not your defaults. They differ by workload. General tasks want temperature 1.0, top_p 0.95, top_k 20, presence_penalty 1.5. Precise coding tasks want temperature 0.6 and presence_penalty 0.0. These are tuned to how the model was trained.
Context
262,144 tokens out of the box. You can push toward roughly 1M with YaRN at a scaling factor of 4.0, supported natively in both vLLM and SGLang.
Read their warning on this one. Open-source runtimes apply YaRN statically, so the same scaling factor hits every request regardless of length, which can hurt quality on ordinary inputs. Their guidance is to enable it only when the workload needs it, and to size the factor to the actual target window rather than maxing it out.
What it plugs into
The model exposes an OpenAI-compatible endpoint with tool calling, and emits well-formed function calls that get parsed into standard tool_calls. That means most agent frameworks work without custom glue.
The card lists Ollama, llama.cpp via llama-server, Hermes Agent, OpenClaw, and Unsloth Studio for fine-tuning. For terminal coding agents there’s a config example for registering your local endpoint as a provider in OpenCode.
Worth watching
The 35B is the size I’d point most people at. Roughly 3B active parameters per token, 79.0 on SWE-bench Verified, and it beats its like-for-like peer across coding and agentic benchmarks.
The bigger question is whether the self-improvement loop keeps paying off. If a model can generate its own curriculum, the bottleneck moves from human task curation to compute. That’s a different scaling story than the one the field has been running on, and the next release will say more about it than this one does.
For now: the weights are up, the license is MIT, and the tables are published with the methodology attached. Go check the numbers yourself.
Sources
- Ornith-1.5: From Self-Scaffolding to Self-Improvement — Ornith Team, Aug. 2026
- ornith-ai/Ornith-1.5-9B model card — Hugging Face
Not sponsored, not affiliated with Ornith. All benchmark figures belong to the Ornith team.