Ornith AI open-sources Ornith-1.5, a self-improving LLM family it says rivals Claude Opus 4.8
Ornith AI released Ornith-1.5, a family of open-weight large language models under the MIT license, on August 18, 2026, claiming its top model matches Claude Opus 4.8 on select benchmarks.
Ornith AI released Ornith-1.5, a family of open-weight large language models it says can match Anthropic’s Claude Opus 4.8 on select agentic benchmarks, under the permissive MIT license on August 18, 2026. The family comes in three sizes: a 9-billion-parameter dense model, a 35-billion-parameter mixture-of-experts model and a 397-billion-parameter mixture-of-experts model.
The release extends what Ornith calls a self-improvement loop. Rather than training on fixed, human-curated tasks, the model proposes new training tasks, builds task-specific scaffolds and generates solution attempts that are optimized together using reinforcement learning, according to Ornith’s technical post. The company says the method targets reasoning, agentic and coding work.
On its published model card, Ornith reports the 9B model scoring 70.6 on SWE-bench Verified and 46.2 on Terminal-Bench 2.1. In its technical post, Ornith says the 397B model reaches 86.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE, which it says matches Claude Opus 4.8 on those tests. The 9B model runs on a single GPU, and a quantized “Mobile” variant runs on phones, the company says. All checkpoints, including FP8, GGUF, MLX and NVFP4 quantizations, are on Hugging Face.
The benchmark claims are Ornith’s own and have not been independently verified. Self-reported scores on agentic benchmarks are sensitive to prompt scaffolding and test harness, and “matches Claude Opus 4.8” rests on two hand-picked tests, not a full comparison.
Open weights under the MIT license are the durable part. Anyone can now run the scores down, which is how a self-improvement claim earns or loses credibility.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

Google Research releases MilleMiglia logistics benchmark generator

AWS launches AgentCore Runtime V2 with elastic memory and snapshot starts
