

Inkling is Thinking Machines’ first open-weights model, a 975B MoE with 41B active parameters, 1M context, native reasoning across text, images, and audio, and controllable thinking effort. Fine-tune it on Tinker or download the Apache 2.0 weights.
Loading comments…
Project Info
Product Keywords
Inkling is Thinking Machines’ first open-weights model, released under the Apache 2.0 license. It is a Mixture-of-Experts (MoE) transformer with 975 billion total parameters and 41 billion active parameters per forward pass. Inkling supports a context window of up to 1 million tokens and was pretrained on 45 trillion tokens spanning text, images, audio, and video. It reasons natively across these modalities and offers controllable thinking effort, allowing users to balance cost and performance. The model is available for fine-tuning on Tinker, and its full weights can be downloaded directly.
Inkling uses a Mixture-of-Experts architecture that activates only a fraction of its total parameters per token. This design delivers the capability of a large model while keeping inference compute efficient and latency manageable.
The model supports up to 1 million tokens of context, making it suitable for long-document analysis, extended conversations, and tasks that require reasoning over large amounts of information in a single pass.
Inkling was pretrained on text, images, audio, and video, and it reasons across these modalities natively. It does not rely on separate encoders or pipelines for different input types.
Users can adjust how much computation the model spends on reasoning before generating an answer. This allows fine-grained control over the trade-off between response quality and cost or latency.
Inkling is not the strongest overall model available today — instead, a combination of qualities makes it a good open-weights base for customization.
This honest positioning sets Inkling apart. Rather than chasing top benchmark scores, Thinking Machines focused on building a balanced, flexible foundation model that excels as a starting point for fine-tuning. The model is available on Tinker for immediate customization, and the team demonstrated this by having Inkling fine-tune itself — writing its own training job, running it, and evaluating the result — all within the Tinker platform.
You are looking for a permissively licensed, open-weights model that balances multimodal capability with efficient inference, and you value the ability to fine-tune and customize the model for your own use cases. Inkling is especially worth exploring if you want to experiment with controllable reasoning effort or need a model that can handle long contexts across text, images, and audio.
Other tools you might consider
Run state-of-the-art open-source models (GLM 5.1, Kimi K2.7 Code, MiniMax M2.7, and more) in Claude Code at up to 4× the speed (up to 200 tok/s) for a flat $29/month. Set up in minutes, no code changes.
The world can't build compute fast enough to keep up with AI demand. So we took a different path. ZeroGPU is AI infrastructure powered by small language models running on a hybrid edge network reusing compute that already exists. Not every task needs a frontier model. Our purpose-built, edge-optimized models run 10x faster, 50% cheaper and offload 70–80% of production tasks to small models with frontier-level accuracy.
The moment an agent needs to deploy something, it slams face-first into a wall built for humans. Today we're rolling out Temporary Accounts on Cloudflare Workers. Any agent can now run wrangler deploy — temporary and get a live Worker in seconds.
GitHits gives coding agents access to the open-source code your app depends on. Get real implementation examples, dependency source navigation, package inspection and documentation. Agents can grep and read your codebase. They can't grep and read the open-source code your app depends on. That's where they start guessing, retrying, and looping. GitHits builds a version-aware index on demand. Agents can search, navigate, and inspect the code behind their dependencies. CLI: npx githits@latest init
Maker
neon_dev
Loading comments…