📊 Full opportunity report: How Inkling Is Changing The Landscape Of AI Technology on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Thinking Machines has launched Inkling, a large-scale multimodal AI model, on Hugging Face. It supports text, images, and audio reasoning but requires substantial computing resources. Its full capabilities and licensing details remain unverified, highlighting the need for further evaluation of such models, as discussed in the original analysis.
Thinking Machines has released Inkling on Hugging Face, presenting it as an open multimodal model capable of processing text, images, and audio within a claimed one-million-token context window. For more details, see the original analysis. The release marks a significant step in making large-scale multimodal AI more accessible, although its demanding hardware requirements limit immediate widespread deployment.
Inkling is a decoder-only Mixture-of-Experts model with a total of 975 billion parameters and 41 billion active during processing, trained on 45 trillion tokens across multiple data modalities, including text, images, audio, and video. The architecture employs 256 experts, utilizing a combination of global and sliding-window attention, with image inputs processed through hierarchical patching and audio converted into mel-spectrograms. This approach reflects recent advances in multimodal AI models, as detailed in the original analysis. The release includes BF16 and NVFP4 checkpoints optimized for different hardware, with the former requiring approximately 2 TB of VRAM and the latter about 600 GB, making full deployment feasible only on high-end systems or via hosted inference services.
Implications of Inkling’s Open Multimodal Approach
Inkling’s release introduces a powerful, open-access multimodal AI that can reason across text, images, and audio, potentially accelerating developments in scientific research, media analysis, and enterprise applications. However, the high hardware demands limit its immediate use to organizations with significant computational resources or those relying on hosted services. This development signals a shift towards larger, more integrated models, but also raises questions about accessibility, safety, and licensing.

PNY VCNRTXA6000-SB NVIDIA RTX A6000 Graphics Card 48GB GDDR6
NVIDIA Virtual PC (vPC)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on Large-Scale Multimodal Models
Prior to Inkling, multimodal AI models have largely been proprietary or limited in scale, with notable examples like GPT-4 and PaLM-E. The release of Inkling by Thinking Machines, with its unprecedented parameter count, follows a trend of increasing model size to improve reasoning and understanding capabilities. The model’s architecture, based on sparse Mixture-of-Experts design, aims to balance scale with efficiency, though practical deployment remains challenging due to hardware constraints. The announcement comes amid a broader industry push for open, versatile models capable of multi-sensory reasoning.
“This model is huge.”
— Hugging Face

Self-Hosted AI Infrastructure: Deploy, Manage, and Scale LLMs on Proxmox, Docker, and NAS (Developer guides)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unverified Aspects of Inkling’s Capabilities and Deployment
There are no independent benchmark results or safety evaluations available at this stage. The actual performance of Inkling on real-world tasks, especially in video processing, remains unconfirmed. Licensing terms, usage restrictions, and the availability of training data or code are also unclear, raising questions about accessibility and safety.

ESP32 Basic Starter Ai Chatbot Kit Development Board USB-C Dual Core Microcontroller Support AP/STA/AP+STA Compatible with Arduino IDE IoT for Beginners Engineers
【High Performance】The ESP32 module is equipped with a dual-core CPU and features a Type-C USB interface, with 44…
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for Testing and Validation of Inkling
Developers and organizations will begin testing Inkling through supported inference engines like Transformers and llama.cpp. Early evaluations will focus on latency, accuracy, and resource consumption across various workloads. Independent benchmarking, safety assessments, and licensing clarifications are expected to follow, providing a clearer picture of its practical utility and limitations.

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What makes Inkling different from other multimodal models?
Inkling features an unprecedented 975-billion-parameter scale and supports text, images, and audio reasoning within a large context window, aiming for versatile, integrated multimodal capabilities.
Can I run Inkling on a personal computer?
Currently, the hardware requirements are extremely high, with estimates of 2 TB of VRAM for BF16 checkpoints, making it impractical for typical consumer systems. Hosted inference or cloud deployment are the most feasible options.
Will Inkling process video directly?
While the architecture supports image inputs with a temporal dimension, native video processing performance has not yet been evaluated or confirmed.
Is Inkling open-source and freely available?
The release describes Inkling as an open model, but licensing details, usage restrictions, and whether training code will be shared remain unclear at this time.
How does Inkling compare to existing large multimodal models?
Independent benchmarks and safety evaluations are pending, so it is not yet clear how Inkling’s performance and safety profile compare with models like GPT-4 or PaLM-E.
Source: ThorstenMeyerAI.com