Together AI brings Thinking Machines Lab’s new model Inkling on day 0

Imported from official source

AI Classified by Officially

Build with Inkling, a new multimodal mixture-of-experts model designed for advanced reasoning, through Together AI’s production inference platform

Jue Wang, Wei Gong, Yineng Zhang, Hiral Jasani

40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production...

Today, Thinking Machines Lab released Inkling, a new multimodal mixture-of-experts model built for token-efficient reasoning, native multimodal understanding, and broad task versatility. Together AI is excited to collaborate with the Thinking Machines Lab team to make Inkling available to developers on our inference platform. 

Inkling accepts text, image, and audio inputs and produces text outputs through a unified decoder architecture. It supports controllable inference effort, allowing developers to adjust how much reasoning the model applies based on the needs of each task. Its post-training also spans a wide range of capabilities, including scientific reasoning, coding, agentic workflows, forecasting, and calibrated prediction.

Under the hood, Inkling introduces several architectural innovations beyond a conventional decoder-only Transformer, including query-conditioned relative attention, short causal convolutions throughout the model, and a mixture-of-experts architecture with a shared expert sink. Together, these components are designed to support strong reasoning and multimodal capabilities while maintaining efficient model execution. Serving it efficiently at scale is nontrivial, and it's exactly the kind of workload Together AI's inference stack is built to optimize, so you get the model's efficiency gains in practice for production inference. On Together AI, Inkling runs with an optimized FlashAttention-4–based attention kernel designed to efficiently support its query-conditioned relative attention mechanism in production.

Congratulations to the Thinking Machines Lab team on the release.

This is an extract. The publication continues at the source.

Read the original at the source: https://www.together.ai/blog/together-ai-brings-thinking-machines-labs-new-model-inkling-on-day-0

Officially imported this from Together AI’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.

Provenance

Organization
Together AI — imported from official source
Official source
https://www.together.ai/blog/rss.xml RSS
Imported
September 20, 2026 19:52
Versions
1 recorded
Identity
https://www.together.ai/blog/together-ai-brings-thinking-machines-labs-new-model-inklin...

Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.