Skip to content
-
Subscribe to our newsletter & never miss our best posts. Subscribe Now!
AI Brainbox logo AI Brainbox logo AI BrainBox

AI Tools | Automation | Workflows

AI Brainbox logo AI Brainbox logo AI BrainBox

AI Tools | Automation | Workflows

  • Home
  • Tutorials
  • About
  • Contact
  • Privacy Policy
  • Home
  • Tutorials
  • About
  • Contact
  • Privacy Policy
Close

Search

Trending Now:
Free AI Voice Generator Claude code for free Free AI tools Local AI models
  • FaceBook
  • X
  • LinkedIn
  • Instagram
  • YouTube
Subscribe
AI Brainbox logo AI Brainbox logo AI BrainBox

AI Tools | Automation | Workflows

AI Brainbox logo AI Brainbox logo AI BrainBox

AI Tools | Automation | Workflows

  • Home
  • Tutorials
  • About
  • Contact
  • Privacy Policy
  • Home
  • Tutorials
  • About
  • Contact
  • Privacy Policy
Close

Search

Trending Now:
Free AI Voice Generator Claude code for free Free AI tools Local AI models
  • FaceBook
  • X
  • LinkedIn
  • Instagram
  • YouTube
Subscribe
Home/AI Models/Ornith-1.5: The Open Model That Beats Claude
AI Models

Ornith-1.5: The Open Model That Beats Claude

By AI BrainBox
August 21, 2026 6 Min Read
0

Most model releases follow the same script. Bigger model, more data, a few points on the leaderboard, a blog post full of adjectives.

Ornith-1.5 is doing something else, and the something else is worth understanding even if you never download the weights.

During training, this model writes its own practice problems. It builds the grading scripts for those problems. It attempts them, and the results feed back into all three of those steps at once. There is no fixed pile of human-written tasks sitting underneath it. The curriculum generates itself and gets harder as the model gets better.

That’s the pitch. Here’s what’s actually in the release, what the numbers show, and where the framing runs ahead of the data.

https://youtu.be/XMFftQi3teo

What shipped

Three sizes, all MIT licensed:

ModelTypeNotes
Ornith-1.5-397BMoEThe flagship
Ornith-1.5-35B-A3BMoERoughly 3B parameters active per token
Ornith-1.5-9BDenseSingle-GPU, with a quantized Mobile variant

All three extend Ornith-1.0, which the team built on top of Qwen3.5 and Gemma 4 with additional continued pretraining, mid-training, and post-training.

The 9B has a quantized sibling called Ornith-1.5-9B-Mobile that the blog says targets iPhone and Android. The blog doesn’t give a file size for it, so I’m not going to invent one.

How the self-improvement loop works

Training a coding model normally means humans write thousands of problems and the test scripts that verify solutions. The model grinds against that fixed set. It works, but you eventually run out of problems, and the ones you wrote may not target what the model is actually bad at.

Ornith-1.5 removes the fixed set. Each training cycle runs three stages.

The model proposes a task. It gets an environment or codebase, high-level guidance about the kind of task to produce, and its own history of what it has already solved. From that it writes a harder problem, aimed at a gap in its own ability.

The model builds a scaffold. In Ornith’s terms, the scaffold is the instructions, the tools, the decomposition strategy, and the orchestration used to approach the problem. Not a prompt template. The whole apparatus around the task.

The model attempts a solution. Conditioned on both the task and the scaffold, the policy produces a rollout.

Then reward from that rollout propagates back through all three stages. The model isn’t only learning to solve better. It’s learning to write more useful tasks and build more effective scaffolds, in the same loop.

Run that repeatedly and stronger policies produce harder tasks, harder tasks produce better training signal, and the scaffolds keep evolving toward whatever actually elicits the model’s capability.

The part that keeps it honest

If a model invents its own homework, the obvious failure mode is that it invents easy homework and farms a great score.

Ornith scores every generated task on three signals and multiplies them together, so a task has to satisfy all three or the reward collapses.

Validity. Is the task coherent and solvable? Does the scaffold execute correctly and evaluate candidates reliably? The checks include whether high-confidence solutions pass and clearly incorrect ones fail. This is a hard gate. If validity comes back zero, the entire task reward is zero, which stops malformed tasks from getting paid just for looking difficult.

Frontier difficulty. They sample multiple rollouts per task and measure the empirical success rate, then reward tasks whose rate sits near a target. That target is set at 0.2. Roughly one attempt in five succeeds, which is hard enough to be worth learning from and easy enough to produce usable successful trajectories. The elegant part: as the model starts reliably solving a task, that task’s reward drops on its own, pushing the generator toward harder ones.

Novelty. Each task is compared against a buffer of previously generated and trained-on tasks. Too similar, and novelty drops. Without this term, frontier difficulty alone would encourage a thousand slight variations of the same problem.

There’s a parallel reward for the harness itself, scored on task alignment, reward fidelity, and hack resistance. All three stages are optimized with GRPO.

The benchmarks

397B

The flagship is the one trading blows at the top.

BenchmarkOrnith-1.5-397BClaude Opus 4.8
Terminal-Bench 2.1 (Terminus-2)86.185.0
SWE-bench Verified8685.8
DeepSWE5659

It takes Terminal-Bench and SWE-bench Verified by narrow margins and loses DeepSWE. “Matches” is the accurate word. “Beats” isn’t.

35B

The 35B is arguably the most interesting one in the family. It activates only about 3B parameters per token and still posts 68.5 on Terminal-Bench 2.1 against Gemma-4-31B’s 43.4 and Muse-Glimmer-30B’s 51.7. On SWE-bench Verified it hits 79.0 against 52.0 and 76.0 respectively.

It also beats Qwen3.6-35B-A3B across the coding and agentic benchmarks, which is a like-for-like comparison at the same scale.

9B

This is the one most people can actually run.

BenchmarkOrnith-1.5-9BOrnith-1.0-9BGemma-4-31BQwen3.6-35B-A3B
Terminal-Bench 2.1 (Terminus-2)46.243.142.152.5
Terminal-Bench 2.1 (Claude Code)4740.6–49.2
SWE-bench Verified70.669.45273.4
SWE-bench Pro47.542.935.749.5
GPQA Diamond86.482.584.386
BrowseComp56.444.8–62
ClawEval66.563.148.568.7

Against Gemma-4-31B the 9B wins clearly, and SWE-bench Verified at 70.6 against 52 is a wide gap in favor of a model less than a third the size.

The BrowseComp jump is the one I’d flag as most notable. 44.8 to 56.4 in a single version is a big move.

Where the framing overshoots

The blog says the 9B substantially outperforms larger models including Gemma-4-31B and Qwen3.6-35B.

The first half holds. The second doesn’t, and their own table shows it. Qwen3.6-35B-A3B beats the 9B on Terminal-Bench (52.5 to 46.2), SWE-bench Verified (73.4 to 70.6), SWE-bench Pro (49.5 to 47.5), BrowseComp (62 to 56.4), and ClawEval (68.7 to 66.5).

The 9B does win on NL2Repo (32.4 to 29.4), SWE Atlas QnA (20.6 to 15.5), and edges GPQA Diamond (86.4 to 86). So it’s a real fight against a model roughly four times its size, which is genuinely impressive. It just isn’t a sweep, and publishing the table that contradicts your own summary sentence is an odd choice.

Worth crediting the methodology, though. Results are averaged over five independent runs. For SWE-bench evaluations they stripped git history out of the repository images so the model couldn’t look up the commit that fixed the issue, and disabled network access so it couldn’t retrieve the answer. Those are the safeguards you’d want to see.

These are still self-reported numbers from the team’s own evaluation setup. Normal for a launch, but worth holding lightly until others reproduce them.

Running the 9B

The model card is specific about requirements. Ornith-1.5-9B is roughly 19 GB in bf16 and serves on a single 80GB GPU.

You’ll need recent runtimes: Transformers 5.8.1+, vLLM 0.19.1+, or SGLang 0.5.9+.

vLLM:

bash

vllm serve ornith-ai/Ornith-1.5-9B \
  --served-model-name Ornith-1.5-9B \
  --host 0.0.0.0 --port 8000 \
  --max-model-len 262144 \
  --gpu-memory-utilization 0.90 \
  --enable-prefix-caching \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml \
  --reasoning-parser qwen3 \
  --trust-remote-code

SGLang:

bash

python -m sglang.launch_server \
  --model-path ornith-ai/Ornith-1.5-9B \
  --served-model-name Ornith-1.5-9B \
  --host 0.0.0.0 --port 8000 \
  --context-length 262144 \
  --mem-fraction-static 0.85 \
  --tool-call-parser qwen3_coder \
  --reasoning-parser qwen3

Ollama:

bash

ollama run ornith-1.5:9b

Two things to know before you start sending requests.

It’s a reasoning model. The assistant turn opens with a thinking block before the final answer. The serving recipes above enable a reasoning parser so that trace comes back in a separate reasoning_content field, which means you can inspect it during development and hide it in production without string-parsing anything.

Use their sampling parameters, not your defaults. They differ by workload. General tasks want temperature 1.0, top_p 0.95, top_k 20, presence_penalty 1.5. Precise coding tasks want temperature 0.6 and presence_penalty 0.0. These are tuned to how the model was trained.

Context

262,144 tokens out of the box. You can push toward roughly 1M with YaRN at a scaling factor of 4.0, supported natively in both vLLM and SGLang.

Read their warning on this one. Open-source runtimes apply YaRN statically, so the same scaling factor hits every request regardless of length, which can hurt quality on ordinary inputs. Their guidance is to enable it only when the workload needs it, and to size the factor to the actual target window rather than maxing it out.

What it plugs into

The model exposes an OpenAI-compatible endpoint with tool calling, and emits well-formed function calls that get parsed into standard tool_calls. That means most agent frameworks work without custom glue.

The card lists Ollama, llama.cpp via llama-server, Hermes Agent, OpenClaw, and Unsloth Studio for fine-tuning. For terminal coding agents there’s a config example for registering your local endpoint as a provider in OpenCode.

Worth watching

The 35B is the size I’d point most people at. Roughly 3B active parameters per token, 79.0 on SWE-bench Verified, and it beats its like-for-like peer across coding and agentic benchmarks.

The bigger question is whether the self-improvement loop keeps paying off. If a model can generate its own curriculum, the bottleneck moves from human task curation to compute. That’s a different scaling story than the one the field has been running on, and the next release will say more about it than this one does.

For now: the weights are up, the license is MIT, and the tables are published with the methodology attached. Go check the numbers yourself.


Sources

  • Ornith-1.5: From Self-Scaffolding to Self-Improvement — Ornith Team, Aug. 2026
  • ornith-ai/Ornith-1.5-9B model card — Hugging Face

Not sponsored, not affiliated with Ornith. All benchmark figures belong to the Ornith team.

Tags:

free ai modelsornith 1.5ornith-1.5
Author

AI BrainBox

Follow Me
Other Articles
local ai from usb pen drive
Previous

Run an AI Model From a Pen Drive (Windows, Fully Offline)

Recent Posts

  • Ornith-1.5: The Open Model That Beats Claude
  • Run an AI Model From a Pen Drive (Windows, Fully Offline)
  • Unsloth Desktop on Windows: the setup guide I wish I’d had
  • Run Claude Code on Free AI Models with OmniRoute (Beginner Setup Guide)
  • DeepSeek V4 Flash 0731 Just Dropped | Run it FREE | It’s Really Insane
Hey, I’m Umair. I’m a Data Scientist passionate about AI, exploring emerging technologies, and sharing practical tutorials that help people work smarter with AI.
  • X
  • Instagram
  • Facebook
  • YouTube
Open to AI Projects
Get In Touch

Recent Posts

  • Ornith-1.5: The Open Model That Beats Claude
    by AI BrainBox
    August 21, 2026
  • clone-voice-free-voicebox-tutorial
    Clone Any Voice for Free with Voicebox: Full 2026 Guide
    by AI BrainBox
    June 26, 2026
  • chatgpt-prompt-engineering
    ChatGPT Prompt Engineering for Beginners: How to Get Better Results Every Time
    by AI BrainBox
    June 27, 2026
  • how-to-run-ai-locally-ollama-guide
    How to Run AI Locally on Your PC with Ollama (No Cloud, No Subscription)
    by AI BrainBox
    June 27, 2026
ai brainbox

We're exploring the latest and greatest AI tools and techniques, providing you with everything you need to know to keep up with this rapidly evolving field. Passionate about making AI accessible and understandable to everyone.

  • Facebook
  • X
  • Instagram
  • LinkedIn

Latest Posts

  • Ornith-1.5: The Open Model That Beats Claude
  • Run an AI Model From a Pen Drive (Windows, Fully Offline)
  • Unsloth Desktop on Windows: the setup guide I wish I’d had
  • Run Claude Code on Free AI Models with OmniRoute (Beginner Setup Guide)
  • DeepSeek V4 Flash 0731 Just Dropped | Run it FREE | It’s Really Insane

Menu

  • Home
  • Tutorials
  • About
  • Contact
  • Privacy Policy

Contact

Email

contact@aibrainbox.io

Location

New York, USA

Copyright 2026 — AI BrainBox. All rights reserved.