📊 Full opportunity report: Unlocking AI Potential With Inkling By Thinking Machines on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Thinking Machines has unveiled Inkling, a large-scale multimodal AI model accessible on Hugging Face. Its impressive size offers new multimodal reasoning capabilities but requires substantial hardware, limiting immediate broad deployment.

Thinking Machines has released Inkling on Hugging Face, offering a 975-billion-parameter multimodal model capable of processing text, images, and audio within a single framework. For more context, see the significance of Thinking Machines’ hints in AI development. This release is notable because it provides open access to one of the largest models of its kind, although its demanding hardware requirements limit immediate use for most developers.

Inkling is described as a decoder-only Mixture-of-Experts model with 975 billion total parameters and 41 billion active during processing, trained on 45 trillion tokens across multiple data modalities, including text, images, audio, and video. It features a hierarchical patching module for image inputs and converts audio into mel-spectrograms for processing. The model is available in BF16 and NVFP4 checkpoints, with the latter requiring approximately 600 GB of VRAM, making it accessible primarily through hosted inference or specialized hardware.

Hugging Face reports support for Inkling in popular frameworks like Transformers, SGLang, vLLM, and llama.cpp. Learn more about the development process in Welcome Inkling by Thinking Machines. However, the release lacks independent benchmark results, safety evaluations, or licensing details, raising questions about its performance and usage restrictions. The architecture employs a sparse Mixture-of-Experts approach, activating only a subset of parameters per input to optimize efficiency. This approach is discussed in detail in the original analysis.

At a glance
announcementWhen: announced July 2026
The developmentThinking Machines has released Inkling, a 975-billion-parameter multimodal AI model, on Hugging Face, marking a significant step in open access to large-scale AI models.
At a glance
announcementWhen: announced on Hugging Face; the source m…
The developmentThinking Machines has made its Inkling multimodal model available through Hugging Face with day-one support from several major inference frameworks.

Implications of Inkling’s Open Multimodal Scale

This release demonstrates a move toward larger, multimodal AI models that are accessible to the research community, potentially accelerating developments in areas like scientific analysis, media processing, and enterprise applications. However, its substantial hardware demands mean only well-resourced organizations can deploy it directly, while many will rely on hosted services.

Overall, Inkling’s open access at this scale could influence future AI model development, but the absence of independent benchmarks and licensing clarity leaves its practical impact uncertain for now.

Amazon

high VRAM GPU for AI training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Large-Scale Multimodal Models

Recent years have seen rapid growth in large language models, culminating in multimodal architectures that combine text, images, and audio processing. Prior models, such as GPT-4 and PaLM-E, have demonstrated multimodal capabilities but often remain proprietary or limited in scale. Inkling’s release on Hugging Face marks a notable step in providing open access to a model of this size and multimodal scope, although comparable models like Meta’s Llama 2 or Google’s PaLM have yet to match its scale.

Training on 45 trillion tokens across multiple data types suggests a significant investment in data diversity, but independent validation and benchmark results are not yet available, making it difficult to assess its relative performance.

“Inkling’s release pushes the boundaries of open multimodal AI, but its hardware demands are a major barrier for widespread adoption.”

— Thorsten Meyer, AI researcher

Amazon

AI inference server hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Aspects of Inkling’s Performance and Licensing

There are no published independent benchmark results, safety evaluations, or detailed licensing terms for Inkling. Its actual performance on real-world multimodal tasks, especially video, remains unconfirmed. The effectiveness of its speculative multi-token prediction layers and its safety profile are also unknown, leaving questions about its readiness for broad deployment.

Amazon

multimodal AI model hardware requirements

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Testing and Validation of Inkling

Developers and research organizations will likely begin testing Inkling through supported inference engines to evaluate its speed, accuracy, and safety. Independent benchmarking, safety assessments, and domain-specific fine-tuning are expected to follow, which will clarify its practical capabilities and limitations. Monitoring these evaluations will be critical to understanding its impact.

Amazon

large-scale AI model hosting solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Inkling?

Inkling is a large-scale, 975-billion-parameter multimodal AI model from Thinking Machines, capable of processing text, images, and audio within a unified framework, released openly on Hugging Face.

Can I run Inkling on a personal computer?

Unlikely. The BF16 checkpoint requires about 2 TB of VRAM, and the NVFP4 version needs approximately 600 GB, making it feasible only on high-end hardware or via hosted inference services.

Does Inkling support video processing?

While the architecture includes temporal dimensions for video, Hugging Face states native video performance has not yet been evaluated, so its video capabilities remain unconfirmed.

What are the licensing terms for Inkling?

The release describes Inkling as an open model but does not specify licensing restrictions, training data details, or whether source code is available, leaving licensing terms unclear.

When will independent benchmarks or safety evaluations be available?

It is not yet clear when or if independent evaluations will be published. Developers are expected to begin testing soon, which may provide preliminary insights into its performance and safety.

Source: ThorstenMeyerAI.com

You May Also Like

The 2028 Model Lab Endgame: How Six Becomes Two, Three, or Twelve

Forecasting the future of Western frontier AI labs by 2028, this analysis explores three possible scenarios—consolidation to two or three labs, or a fragmented landscape of twelve.

Wi-Wi is wireless time sync at 1 nanosecond

Wi-Wi, a Japanese-developed wireless synchronization protocol, demonstrates time accuracy down to 1 nanosecond, promising significant advances in broadcast and industrial tech.

How Apple Is Using Technology Operations To Combat Trade Secret Espionage

Apple is deploying advanced technology operations to detect and prevent trade secret theft, including legal actions against ex-employees and monitoring tools.

Nativ: Run Frontier Open Models Locally On Your Mac

Nativ launches a tool allowing users to run frontier open models locally on Macs, enhancing privacy and performance for AI developers and enthusiasts.