Thinking Machines Inkling Explained 2026: Murati's First Open-Weight 975B MoE (Apache 2.0)
Ex-OpenAI CTO Mira Murati's Thinking Machines Lab shipped its first model, Inkling: a 975B MoE (41B active) trained on 45 trillion multimodal tokens, released under Apache 2.0. It's positioned not as a leaderboard-topping assistant but as a foundation companies customize, via Hugging Face weights and fine-tuning on Tinker.
Ex-OpenAI CTO Mira Murati’s Thinking Machines Lab has shipped its first model, Inkling (July 15). The direction is the interesting part. Instead of fighting for the top of a leaderboard as the smartest general model, it’s positioned as a foundation companies customize and make their own — and it’s released as Apache 2.0 open weights.
What stands out?
- Scale: 975B total parameters (MoE), about 41B active per task
- Training: 45 trillion tokens — multimodal across text, image, audio and video
- License: Apache 2.0 open weights — free for commercial use
- Distribution: weights on Hugging Face plus fine-tuning live on the company’s Tinker platform

It follows the “build it big, activate a slice” MoE trend while betting on openness. The same open-weight race shows in Alibaba Qwen3.8-Max and DeepSeek V4 Flash.
Why the different direction?
Murati’s team makes a clear bet: what many enterprises want isn’t the top general-purpose AI on a leaderboard, but a model they can genuinely make their own. So rather than a finished assistant like Claude Opus 5, Inkling is positioned as a base to build your own AI on top of.

How do you use it?
| Path | What |
|---|---|
| Hugging Face | Pull the weights and self-host / deploy |
| Tinker | Fine-tune on the company’s platform for your domain and use case |
| Multimodal base | A foundation for systems handling text, image, audio and video |

Frequently Asked Questions
Q. Is Inkling a ready-to-use assistant like ChatGPT? Not quite. Rather than a finished chatbot, it’s positioned as a foundation model that companies fine-tune to build their own AI.
Q. Can I use it commercially? Yes. The weights are Apache 2.0, so commercial use is free. Pull them from Hugging Face to self-host, or fine-tune on Tinker.
Q. 975B — isn’t that heavy? It’s a mixture-of-experts, so each task activates only about 41B rather than the whole model, keeping a very large model relatively efficient to run.
Related posts
New TechHome Humanoid Robots in 2026: Only One You Can Actually Buy (1X NEO at $20K)
This is the robot-vacuum-generation-one moment — and you're standing in it
New TechThe 2026 AI Laptop Guide: What to Buy — Copilot+ PC vs MacBook M5
Don't fall for the TOPS number — memory bandwidth decides real AI speed
New TechGalaxy Z Fold8 & Flip8: Unpacked 2026 Lineup, Prices & Specs
Foldables now have an 'Ultra' — Unpacked 2026 in 3 minutes