Thinking Machines Lab releases Inkling, its first open-weight multimodal model
Thinking Machines Lab, founded by former OpenAI CTO Mira Murati, released Inkling on July 15, 2026, a 975-billion-parameter open-weight multimodal Mixture-of-Experts model.
Thinking Machines Lab, the startup founded by former OpenAI chief technology officer Mira Murati, released Inkling on July 15, 2026, its first open-weight model. Inkling is a 975-billion-parameter Mixture-of-Experts transformer with 41 billion active parameters, trained on 45 trillion tokens of text, image, audio and video.
The release is a direct bet against one-size-fits-all frontier models. Rather than chase the top of every leaderboard, Thinking Machines is pitching Inkling as a token-efficient, openly licensed model that developers can run and adapt themselves.
Full weights are published on Hugging Face in BF16 and NVFP4 checkpoints, and inference is available through Together AI, Fireworks, Modal, Databricks and Baseten, plus natively on the company’s own Tinker platform. Thinking Machines also previewed a lighter Inkling-Small variant with 12 billion active parameters.
The company reports 97.1% on the AIME 2026 math benchmark and says Inkling uses roughly a third as many tokens as Nvidia’s Nemotron 3 Ultra for equivalent coding performance on one benchmark. Those are the company’s own numbers and have not been independently verified. Thinking Machines said plainly that Inkling is “not the strongest overall model available today.”
Open weights make the token-efficiency claim checkable in a way closed models are not. Whether independent evaluations reproduce the AIME score and the coding-cost edge will determine if Inkling earns a place next to the closed frontier.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

Google Research releases MilleMiglia logistics benchmark generator

AWS launches AgentCore Runtime V2 with elastic memory and snapshot starts
