How Inkling Is Changing The Landscape Of AI Technology

📊 Full opportunity report: How Inkling Is Changing The Landscape Of AI Technology on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Thinking Machines has launched Inkling, a large-scale multimodal AI model, on Hugging Face. It supports text, images, and audio reasoning but requires substantial computing resources. Its full capabilities and licensing details remain unverified, highlighting the need for further evaluation of such models, as discussed in the original analysis.

Thinking Machines has released Inkling on Hugging Face, presenting it as an open multimodal model capable of processing text, images, and audio within a claimed one-million-token context window. For more details, see the original analysis. The release marks a significant step in making large-scale multimodal AI more accessible, although its demanding hardware requirements limit immediate widespread deployment.

Inkling is a decoder-only Mixture-of-Experts model with a total of 975 billion parameters and 41 billion active during processing, trained on 45 trillion tokens across multiple data modalities, including text, images, audio, and video. The architecture employs 256 experts, utilizing a combination of global and sliding-window attention, with image inputs processed through hierarchical patching and audio converted into mel-spectrograms. This approach reflects recent advances in multimodal AI models, as detailed in the original analysis. The release includes BF16 and NVFP4 checkpoints optimized for different hardware, with the former requiring approximately 2 TB of VRAM and the latter about 600 GB, making full deployment feasible only on high-end systems or via hosted inference services.

At a glance
announcementWhen: announced July 2026
The developmentThinking Machines released Inkling, a 975-billion-parameter multimodal model, on Hugging Face, enabling advanced AI applications across multiple data types.
At a glance
announcementWhen: announced on Hugging Face; the source m…
The developmentThinking Machines has made its Inkling multimodal model available through Hugging Face with day-one support from several major inference frameworks.

Implications of Inkling’s Open Multimodal Approach

Inkling’s release introduces a powerful, open-access multimodal AI that can reason across text, images, and audio, potentially accelerating developments in scientific research, media analysis, and enterprise applications. However, the high hardware demands limit its immediate use to organizations with significant computational resources or those relying on hosted services. This development signals a shift towards larger, more integrated models, but also raises questions about accessibility, safety, and licensing.

NVIDIA Tesla V100 Volta GPU Accelerator 32GB Graphics Card

NVIDIA Tesla V100 Volta GPU Accelerator 32GB Graphics Card

  • Interface: PCIe

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Large-Scale Multimodal Models

Prior to Inkling, multimodal AI models have largely been proprietary or limited in scale, with notable examples like GPT-4 and PaLM-E. The release of Inkling by Thinking Machines, with its unprecedented parameter count, follows a trend of increasing model size to improve reasoning and understanding capabilities. The model’s architecture, based on sparse Mixture-of-Experts design, aims to balance scale with efficiency, though practical deployment remains challenging due to hardware constraints. The announcement comes amid a broader industry push for open, versatile models capable of multi-sensory reasoning.

“This model is huge.”

— Hugging Face

Self-Hosted AI Infrastructure: Deploy, Manage, and Scale LLMs on Proxmox, Docker, and NAS (Developer guides)

Self-Hosted AI Infrastructure: Deploy, Manage, and Scale LLMs on Proxmox, Docker, and NAS (Developer guides)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Aspects of Inkling’s Capabilities and Deployment

There are no independent benchmark results or safety evaluations available at this stage. The actual performance of Inkling on real-world tasks, especially in video processing, remains unconfirmed. Licensing terms, usage restrictions, and the availability of training data or code are also unclear, raising questions about accessibility and safety.

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

  • Powerful Microcontroller: ESP32-S3 with 16MB Flash and 8MB PSRAM
  • AI Vision & Voice Capabilities: Camera and audio for AI interactions
  • Multiple Development Platforms: Supports Arduino IDE and ESP-IDF

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Testing and Validation of Inkling

Developers and organizations will begin testing Inkling through supported inference engines like Transformers and llama.cpp. Early evaluations will focus on latency, accuracy, and resource consumption across various workloads. Independent benchmarking, safety assessments, and licensing clarifications are expected to follow, providing a clearer picture of its practical utility and limitations.

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What makes Inkling different from other multimodal models?

Inkling features an unprecedented 975-billion-parameter scale and supports text, images, and audio reasoning within a large context window, aiming for versatile, integrated multimodal capabilities.

Can I run Inkling on a personal computer?

Currently, the hardware requirements are extremely high, with estimates of 2 TB of VRAM for BF16 checkpoints, making it impractical for typical consumer systems. Hosted inference or cloud deployment are the most feasible options.

Will Inkling process video directly?

While the architecture supports image inputs with a temporal dimension, native video processing performance has not yet been evaluated or confirmed.

Is Inkling open-source and freely available?

The release describes Inkling as an open model, but licensing details, usage restrictions, and whether training code will be shared remain unclear at this time.

How does Inkling compare to existing large multimodal models?

Independent benchmarks and safety evaluations are pending, so it is not yet clear how Inkling’s performance and safety profile compare with models like GPT-4 or PaLM-E.

Source: ThorstenMeyerAI.com

You May Also Like

Grok 4.6: SpaceXAI’s New AI Model Eyeing Dominance Over GPT-5.6 & Fable 5

SpaceXAI unveils Grok 4.6, aiming to rival GPT-5.6 and Fable 5 in coding and autonomous tasks, with claimed performance gains and lower costs.

OlmoEarth’s AI Platform: The Next Step In Earth Observation Technology

Ai2 unveils OlmoEarth, a new AI platform claiming to process continent-scale satellite data in about a day, transforming Earth observation.

TypeScript 7

TypeScript 7 has been officially announced, bringing significant updates to the popular programming language. Details are emerging about its new capabilities.

Young Immune Cell Therapy Reverses Cognitive Decline in Mice

Incredible advances suggest young immune cell therapy may reverse cognitive decline by enhancing brain health, but the full potential and implications remain to be explored.