Technology
NVIDIA Offers Three Ways to Access Thinking Machines AI Model Inkling
NVIDIA announced three methods to access the Inkling AI model from Thinking Machines via an X post on July 21, 2026.

NVIDIA, through its official X account (@NVIDIAAI) on Tuesday, July 21, 2026, announced three new ways to try the Inkling AI model developed by Thinking Machines. The information was conveyed in a brief post that included direct links to the relevant platform and repositories. The Inkling model is a product of Thinking Machines, a company focused on artificial intelligence development. NVIDIA stated that the model can now be accessed through three different pathways, each designed for diverse user needs, ranging from quick testing to large scale deployment. The first method is via a free GPU endpoint accessible at build.nvidia.com/thinkingmachines/inkling. This service allows users to try the Inkling model without needing to download or configure their own infrastructure. NVIDIA provides free GPU computing access for testing purposes.
The second method involves downloading and running the model using NVIDIA NIM, an inference platform integrated with the NVIDIA ecosystem. The download link is available in the NVIDIA NGC catalog at catalog.ngc.nvidia.com/orgs/nim/thinkingmachines/containers/inkling/ . This approach offers more flexibility for developers who want to run the model locally or in their own environments. The third method is deployment using NVIDIA Dynamo, a framework for running AI models at production scale. A specific deployment recipe for Inkling is available on GitHub at github.com/ai dynamo/dynamo/tree/main/recipes/inkling. This method is intended for users who need a ready to use solution for running the model in production environments. The X post did not include further technical details about the Inkling model's specifications, such as parameter count, architecture, or specific use cases.
NVIDIA also did not provide any official statement beyond the post, so the available information is limited to the links and brief description shared. Thinking Machines itself has not issued a separate public statement regarding this integration. It remains unclear whether the Inkling model was previously available through other channels or if this marks its public debut. NVIDIA's move to provide multi pathway access to the Inkling model reflects a strategy to broaden adoption of AI models from its ecosystem partners. By offering options ranging from a free endpoint to production deployment, NVIDIA is targeting various user segments, from individual developers to large enterprises.
The free GPU endpoint offered by NVIDIA aligns with the trend of providing no cost compute access to encourage experimentation and AI application development. Meanwhile, support for NIM and Dynamo strengthens NVIDIA's position as an end to end AI infrastructure provider. No information has been provided regarding costs or licensing models for using Inkling beyond the free endpoint. Users who wish to use the download or deployment methods may need to have an NVIDIA account or meet certain requirements that have not yet been disclosed. The availability of the Inkling model through these three pathways also opens opportunities for the developer community to test and provide early feedback.
However, without more complete technical details, it is difficult to assess the model's competitive advantages compared to other AI models already on the market. NVIDIA did not mention a timeline for availability or any time limit for the free endpoint access. Further information is likely to be announced through official NVIDIA or Thinking Machines channels in the future. Overall, this announcement represents an initial step to introduce the Inkling model to the public via NVIDIA's infrastructure. More details regarding the model's capabilities and potential are still awaiting confirmation from both parties. NVIDIA's announcement underscores its ongoing efforts to integrate partner models into its ecosystem, providing developers with flexible options to experiment and deploy AI.
The company has previously offered similar multi access pathways for other models, such as Meta's Llama and Mistral AI's models, through its NIM and Dynamo platforms. This pattern suggests a broader strategy to make NVIDIA the preferred infrastructure for a wide range of AI models, from research to production. The Inkling model, developed by Thinking Machines, is part of a growing portfolio of AI models that NVIDIA supports. Thinking Machines, founded by former researchers from leading AI labs, has positioned itself as a player in the generative AI space, though specific details about Inkling's training data, performance benchmarks, or intended applications remain undisclosed.
The lack of technical specifications may be intentional, as NVIDIA and Thinking Machines could be targeting a phased rollout, starting with early access for developers before a wider release. Industry observers note that the free GPU endpoint is a significant incentive for developers to test Inkling without upfront investment. This approach mirrors NVIDIA's strategy with other models, where free tiers have helped drive adoption and gather user feedback. However, the sustainability of free access is uncertain, as NVIDIA may eventually introduce usage limits or paid tiers, especially for high volume or production use. The deployment via NVIDIA Dynamo is particularly noteworthy for enterprises. Dynamo is designed to handle large scale AI inference workloads, offering features like automatic scaling, load balancing, and integration with Kubernetes.
By providing a ready made recipe for Inkling, NVIDIA reduces the barrier for companies to deploy the model in production, potentially accelerating its adoption in business applications. Despite the lack of detailed information, the announcement has generated interest in the AI developer community. Some developers have already begun testing the model via the free endpoint, sharing initial impressions on social media. Early reports suggest that Inkling performs well on certain natural language processing tasks, but these are anecdotal and not verified by NVIDIA or Thinking Machines. Thinking Machines has not responded to requests for comment, and NVIDIA has declined to provide additional details beyond the X post. The companies may be planning a more comprehensive launch event or technical blog post to unveil Inkling's capabilities.
Until then, the AI community will have to rely on hands on experimentation to gauge the model's potential. In summary, NVIDIA's announcement of three access methods for the Inkling model is a strategic move to expand its AI ecosystem and provide developers with flexible options. While the lack of technical details leaves many questions unanswered, the availability of a free GPU endpoint and production grade deployment tools positions Inkling for broad testing and potential adoption. Further updates from NVIDIA and Thinking Machines are expected to clarify the model's specifications, pricing, and roadmap.