NVIDIA Build Offers Free 1M-Context Kimi K3 and DeepSeek API with Rate Limits
NVIDIA Build now provides free, unmetered API access to open models including Kimi K3, DeepSeek V4.1 Flash, and GLM 5.3 with 1M token contexts. While token quotas are gone, access is capped at 40 requests per minute and requires SMS verification. Developers can point Cursor or agent frameworks to NVIDIA's OpenAI-compatible endpoint immediately.

Impact: Medium
Why it matters
You can route prototyping and agentic coding workflows to 1M-context frontier open models for zero token cost by swapping your base URL.
TL;DR
- 01NVIDIA Build offers unlimited daily API calls to Kimi K3, DeepSeek V4.1 Flash, and GLM 5.3 capped at ~40 RPM.
- 02All four flagship open models provide 1,048,576-token context windows over OpenAI-compatible endpoints.
- 03Account activation requires SMS phone verification, and Kimi K3 requires preserving reasoning_content in multi-turn sessions.
Key facts
- Context Window
- 1,048,576 tokens
- Rate Limit
- ~40 requests per minute (~57,600/day)
- API Pricing
- Free (no credit card required)
- Verification
- SMS phone verification required
Unmetered Access with a 40 RPM Throttle
NVIDIA Build has updated its developer tier, removing the previous 1,000-credit ceiling in favor of unlimited daily requests on hosted endpoints at integrate.api.nvidia.com/v1. Access is strictly rate-limited to approximately 40 requests per minute per API key. For personal agent loops, evaluation pipelines, and development testbeds, 40 RPM translates to a theoretical ceiling of 57,600 requests per day without incurring any cloud compute bills.
Flagship Mixture-of-Experts with 1M Context Windows
The free tier provides access to four prominent open-weight models: Kimi K3, DeepSeek V4.1 Flash, GLM 5.3, and GLM 5.3 Flash. All four models run on mixture-of-experts architectures and provide full 1,048,576-token context windows. Kimi K3 targets long-horizon agentic coding with thinking permanently enabled, while DeepSeek V4.1 Flash and GLM 5.3 Flash offer high-speed inference for lightweight structured generation and multi-modal tool use.
The Verification Barrier and Agent State Retention
Access requires an NVIDIA Developer Program account and an SMS verification code from a physical carrier. Virtual numbers and borrowed SIMs violate NVIDIA terms and risk account revocation, while certain country prefixes currently fail or remain omitted from the portal. Additionally, developers integrating Kimi K3 into agent runtimes must preserve state accurately: because thinking cannot be disabled, multi-turn tool calling requests must pass back the assistant message in full, retaining both reasoning_content and tool_calls payloads to prevent context desynchronization.
Try it in 2 minutes
curl https://integrate.api.nvidia.com/v1/chat/completions -H "Authorization: Bearer $NVIDIA_API_KEY" -H "Content-Type: application/json" -d '{"model": "moonshotai/kimi-k3", "messages": [{"role": "user", "content": "Explain Swift async/await in three sentences"}]}'bash
✓ When to use
- Prototyping long-context agentic coding workflows without incurring token bills.
- Evaluating Kimi K3 or DeepSeek V4.1 Flash on custom benchmarks and 1M context tasks.
✕ When NOT to use
- High-throughput production systems requiring guaranteed SLAs beyond 40 RPM.
- Automated deployment accounts originating from unsupported phone country codes.
What to do today
- Register for an NVIDIA Developer Program account and complete SMS verification.
- Point your coding agent or Cursor custom model base URL to integrate.api.nvidia.com/v1.
- Ensure agent loop headers preserve reasoning_content and tool_calls payloads for Kimi K3.
Sources