Nvidia is giving away powerful AI models and free hosted inference to developers, betting that free software pulls buyers toward its GPU hardware. The company's newest release, Nemotron 3.5 Lightning, runs on a single GPU and supports context windows up to 1 million tokens.
Nvidia launched Nemotron 3.5 Lightning on August 11, 2026, a 30-billion-parameter model built on a Mixture of Experts architecture with roughly 3 billion active parameters. The design lets the model run on a single GPU while supporting context windows up to 1 million tokens, and developers can download it from Hugging Face without paying for it.
Free models, paid hardware
The strategy mirrors a razor-and-blades approach: give away the software that draws developers in, then sell the hardware needed to run it well. Nvidia now provides free hosted inference for over 100 AI models through its build.nvidia.com APIs, letting developers test and build on top of them without a credit card.
Earlier in 2026, Nvidia had already released Nemotron 3 Ultra, a 550-billion-parameter model, alongside Dynamo 1.0, software that reportedly enhances GPU performance by up to 7x.
An open-weight camp forms
The AI industry has split into two camps. OpenAI and Anthropic keep their models closed and proprietary behind API paywalls, while Meta, with its Llama series, and now Nvidia push open-weight alternatives that anyone can download, modify, and deploy.
Nvidia releases the models under permissive licenses covering entire families, including Nemotron and Cosmos. They are optimized for Nvidia-specific hardware formats such as NVFP4, a quantization format built for Nvidia chips. In August 2026, Nvidia also introduced NeMo Switchyard, a tool that routes tasks across different models to reduce inference costs.
Source: Crypto Briefing
Trading involves risk.