Skip to content
HomeNewsDigestsConceptsGuidesToolbox
AboutSubscribeUA
Subscribe

AI Today Brief

The daily AI-engineering brief. Built in public. EN · UA.

XTelegramLinkedInYouTubeRSS

Follow AI Today Brief on LinkedIn for daily AI-engineering updates and the weekly “5 shifts that changed how developers work” PDF.

Explore

NewsDigestsConceptsGuides

Company

SubscribeAdvertiseAbout

Legal

Editorial policyAI disclosurePrivacyTerms

© 2026 AI Today Brief. All rights reserved.

  1. Home
  2. /News
  3. /Tools & releases
  4. /NVIDIA Build Offers Free 1M-Context Kimi K3 and DeepSeek API with Rate Limits
Tools & releases

NVIDIA Build Offers Free 1M-Context Kimi K3 and DeepSeek API with Rate Limits

NVIDIA Build now provides free, unmetered API access to open models including Kimi K3, DeepSeek V4.1 Flash, and GLM 5.3 with 1M token contexts. While token quotas are gone, access is capped at 40 requests per minute and requires SMS verification. Developers can point Cursor or agent frameworks to NVIDIA's OpenAI-compatible endpoint immediately.

October 1, 2026· 7 min read
OKCurated by Oleksandr Kuzmenko, AI Product Engineer·Updated October 1, 2026·Sources cited on every story
AI-assisted · editor-reviewed·How we use AI
NVIDIA Build Offers Free 1M-Context Kimi K3 and DeepSeek API with Rate Limits

Impact: Medium

Why it matters

You can route prototyping and agentic coding workflows to 1M-context frontier open models for zero token cost by swapping your base URL.

TL;DR

  • 01NVIDIA Build offers unlimited daily API calls to Kimi K3, DeepSeek V4.1 Flash, and GLM 5.3 capped at ~40 RPM.
  • 02All four flagship open models provide 1,048,576-token context windows over OpenAI-compatible endpoints.
  • 03Account activation requires SMS phone verification, and Kimi K3 requires preserving reasoning_content in multi-turn sessions.

Key facts

Context Window
1,048,576 tokens
Rate Limit
~40 requests per minute (~57,600/day)
API Pricing
Free (no credit card required)
Verification
SMS phone verification required

Unmetered Access with a 40 RPM Throttle

NVIDIA Build has updated its developer tier, removing the previous 1,000-credit ceiling in favor of unlimited daily requests on hosted endpoints at integrate.api.nvidia.com/v1. Access is strictly rate-limited to approximately 40 requests per minute per API key. For personal agent loops, evaluation pipelines, and development testbeds, 40 RPM translates to a theoretical ceiling of 57,600 requests per day without incurring any cloud compute bills.

Flagship Mixture-of-Experts with 1M Context Windows

The free tier provides access to four prominent open-weight models: Kimi K3, DeepSeek V4.1 Flash, GLM 5.3, and GLM 5.3 Flash. All four models run on mixture-of-experts architectures and provide full 1,048,576-token context windows. Kimi K3 targets long-horizon agentic coding with thinking permanently enabled, while DeepSeek V4.1 Flash and GLM 5.3 Flash offer high-speed inference for lightweight structured generation and multi-modal tool use.

The Verification Barrier and Agent State Retention

Access requires an NVIDIA Developer Program account and an SMS verification code from a physical carrier. Virtual numbers and borrowed SIMs violate NVIDIA terms and risk account revocation, while certain country prefixes currently fail or remain omitted from the portal. Additionally, developers integrating Kimi K3 into agent runtimes must preserve state accurately: because thinking cannot be disabled, multi-turn tool calling requests must pass back the assistant message in full, retaining both reasoning_content and tool_calls payloads to prevent context desynchronization.

Try it in 2 minutes

curl https://integrate.api.nvidia.com/v1/chat/completions -H "Authorization: Bearer $NVIDIA_API_KEY" -H "Content-Type: application/json" -d '{"model": "moonshotai/kimi-k3", "messages": [{"role": "user", "content": "Explain Swift async/await in three sentences"}]}'

bash

✓ When to use

  • Prototyping long-context agentic coding workflows without incurring token bills.
  • Evaluating Kimi K3 or DeepSeek V4.1 Flash on custom benchmarks and 1M context tasks.

✕ When NOT to use

  • High-throughput production systems requiring guaranteed SLAs beyond 40 RPM.
  • Automated deployment accounts originating from unsupported phone country codes.

What to do today

  • →Register for an NVIDIA Developer Program account and complete SMS verification.
  • →Point your coding agent or Cursor custom model base URL to integrate.api.nvidia.com/v1.
  • →Ensure agent loop headers preserve reasoning_content and tool_calls payloads for Kimi K3.
#NVIDIA Build#Kimi K3#DeepSeek V4.1 Flash#GLM 5.3#Cursor

Sources

  • Kimi K3, DeepSeek and GLM free from NVIDIA: I tested the claim and here is the catch
ShareShare on XShare on LinkedIn
← Previous storyGemini 4 Argon Delivers 1M Output Tokens and Deep Code MigrationsNext story →OpenAI Reports Sandboxed Agent Bypassed Proxy via Recursive DNS Tunneling

Related stories

  • Tools & releasesMagnitude Ships On-Device Inference Engine for AI Coding Agents
  • Tools & releasesOpenAI Launches Space and Pages for Collaborative Autonomous Agent Workspaces
  • Tools & releasesAnthropic Evaluates GLM-5.3: Open-Weight Model Matches Frontier Autonomous Cyber Exploit Capabilities
  • Tools & releasesSquint Brings Native In-Place LLM Queries to Any macOS Application

Email digest

Get the morning AI brief

One email a day — the stories that matter for engineers, founders and tech leads. Human-edited, with links to primary sources.

  • ✓120+ sources scanned daily
  • ✓Edited by a human
  • ✓1 email per day
  • ✓EN + UA

By subscribing you agree to the privacy policy.