How Inkling Is Changing The Landscape Of AI Technology

📊 Full opportunity report: How Inkling Is Changing The Landscape Of AI Technology on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Thinking Machines has launched Inkling, a large-scale multimodal AI model, on Hugging Face. It supports text, images, and audio reasoning but requires substantial computing resources. Its full capabilities and licensing details remain unverified, highlighting the need for further evaluation of such models, as discussed in the original analysis.

Thinking Machines has released Inkling on Hugging Face, presenting it as an open multimodal model capable of processing text, images, and audio within a claimed one-million-token context window. For more details, see the original analysis. The release marks a significant step in making large-scale multimodal AI more accessible, although its demanding hardware requirements limit immediate widespread deployment.

Inkling is a decoder-only Mixture-of-Experts model with a total of 975 billion parameters and 41 billion active during processing, trained on 45 trillion tokens across multiple data modalities, including text, images, audio, and video. The architecture employs 256 experts, utilizing a combination of global and sliding-window attention, with image inputs processed through hierarchical patching and audio converted into mel-spectrograms. This approach reflects recent advances in multimodal AI models, as detailed in the original analysis. The release includes BF16 and NVFP4 checkpoints optimized for different hardware, with the former requiring approximately 2 TB of VRAM and the latter about 600 GB, making full deployment feasible only on high-end systems or via hosted inference services.

At a glance
announcementWhen: announced July 2026
The developmentThinking Machines released Inkling, a 975-billion-parameter multimodal model, on Hugging Face, enabling advanced AI applications across multiple data types.
At a glance
announcementWhen: announced on Hugging Face; the source m…
The developmentThinking Machines has made its Inkling multimodal model available through Hugging Face with day-one support from several major inference frameworks.

Implications of Inkling’s Open Multimodal Approach

Inkling’s release introduces a powerful, open-access multimodal AI that can reason across text, images, and audio, potentially accelerating developments in scientific research, media analysis, and enterprise applications. However, the high hardware demands limit its immediate use to organizations with significant computational resources or those relying on hosted services. This development signals a shift towards larger, more integrated models, but also raises questions about accessibility, safety, and licensing.

PNY VCNRTXA6000-SB NVIDIA RTX A6000 Graphics Card 48GB GDDR6

PNY VCNRTXA6000-SB NVIDIA RTX A6000 Graphics Card 48GB GDDR6

NVIDIA Virtual PC (vPC)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Large-Scale Multimodal Models

Prior to Inkling, multimodal AI models have largely been proprietary or limited in scale, with notable examples like GPT-4 and PaLM-E. The release of Inkling by Thinking Machines, with its unprecedented parameter count, follows a trend of increasing model size to improve reasoning and understanding capabilities. The model’s architecture, based on sparse Mixture-of-Experts design, aims to balance scale with efficiency, though practical deployment remains challenging due to hardware constraints. The announcement comes amid a broader industry push for open, versatile models capable of multi-sensory reasoning.

“This model is huge.”

— Hugging Face

Self-Hosted AI Infrastructure: Deploy, Manage, and Scale LLMs on Proxmox, Docker, and NAS (Developer guides)

Self-Hosted AI Infrastructure: Deploy, Manage, and Scale LLMs on Proxmox, Docker, and NAS (Developer guides)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Aspects of Inkling’s Capabilities and Deployment

There are no independent benchmark results or safety evaluations available at this stage. The actual performance of Inkling on real-world tasks, especially in video processing, remains unconfirmed. Licensing terms, usage restrictions, and the availability of training data or code are also unclear, raising questions about accessibility and safety.

ESP32 Basic Starter Ai Chatbot Kit Development Board USB-C Dual Core Microcontroller Support AP/STA/AP+STA Compatible with Arduino IDE IoT for Beginners Engineers

ESP32 Basic Starter Ai Chatbot Kit Development Board USB-C Dual Core Microcontroller Support AP/STA/AP+STA Compatible with Arduino IDE IoT for Beginners Engineers

【High Performance】The ESP32 module is equipped with a dual-core CPU and features a Type-C USB interface, with 44…

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Testing and Validation of Inkling

Developers and organizations will begin testing Inkling through supported inference engines like Transformers and llama.cpp. Early evaluations will focus on latency, accuracy, and resource consumption across various workloads. Independent benchmarking, safety assessments, and licensing clarifications are expected to follow, providing a clearer picture of its practical utility and limitations.

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What makes Inkling different from other multimodal models?

Inkling features an unprecedented 975-billion-parameter scale and supports text, images, and audio reasoning within a large context window, aiming for versatile, integrated multimodal capabilities.

Can I run Inkling on a personal computer?

Currently, the hardware requirements are extremely high, with estimates of 2 TB of VRAM for BF16 checkpoints, making it impractical for typical consumer systems. Hosted inference or cloud deployment are the most feasible options.

Will Inkling process video directly?

While the architecture supports image inputs with a temporal dimension, native video processing performance has not yet been evaluated or confirmed.

Is Inkling open-source and freely available?

The release describes Inkling as an open model, but licensing details, usage restrictions, and whether training code will be shared remain unclear at this time.

How does Inkling compare to existing large multimodal models?

Independent benchmarks and safety evaluations are pending, so it is not yet clear how Inkling’s performance and safety profile compare with models like GPT-4 or PaLM-E.

Source: ThorstenMeyerAI.com

You May Also Like

Show HN: Shirei, Cross-platform GUI Framework In Native Go

Shirei, a new cross-platform GUI framework built in native Go, was showcased on Show HN, aiming to simplify GUI development with native performance.

Incremental – A Library For Incremental Computations

Incremental, a new library for incremental computations, aims to optimize performance in data processing and software development.

What Audio Interfaces Do That USB Mics Never Quite Can

Many audio interfaces offer superior sound quality and control that USB mics can’t match, but there’s more to discover about their true advantages.

Young Immune Cell Therapy Reverses Cognitive Decline in Mice

Incredible advances suggest young immune cell therapy may reverse cognitive decline by enhancing brain health, but the full potential and implications remain to be explored.