Skip to content
HomeNewsDigestsConceptsGuidesToolbox
AboutSubscribeUA
Subscribe

AI Today Brief

The daily AI-engineering brief. Built in public. EN · UA.

XTelegramLinkedInYouTubeRSS

Follow AI Today Brief on LinkedIn for daily AI-engineering updates and the weekly “5 shifts that changed how developers work” PDF.

Explore

NewsDigestsConceptsGuides

Company

SubscribeAdvertiseAbout

Legal

Editorial policyAI disclosurePrivacyTerms

© 2026 AI Today Brief. All rights reserved.

  1. Home/
  2. News/
  3. Token & cost optimization/
  4. Fireworks AI Releases Ember-1: About 40% Fewer Tokens Than Kimi K3
Token & cost optimization

Fireworks AI Releases Ember-1: About 40% Fewer Tokens Than Kimi K3

Fireworks AI has released Ember-1, a model from Fireworks Research built by post-training Moonshot AI's open-weight Kimi K3 to produce shorter reasoning traces while keeping task accuracy. Fireworks reports about 40% fewer tokens across its own evaluations, with Kimi K3's reasoning shortened by 35% to 50% without sacrificing accuracy across seven benchmarks and two customers' production traffic. In one published production A/B run, reasoning tokens dropped 71.3% and total tokens dropped 39% at essentially unchanged task scores.

September 28, 2026· 2 min read
OKCurated by Oleksandr Kuzmenko, AI Product Engineer·Updated September 28, 2026·Sources cited on every story
AI-assisted · editor-reviewed·How we use AI
Fireworks AI Releases Ember-1: About 40% Fewer Tokens Than Kimi K3

Why it matters

Fireworks AI has released Ember-1, a model from Fireworks Research built by post-training Moonshot AI's open-weight Kimi K3 to produce shorter reasoning traces while keeping task accuracy. Fireworks reports about 40% fewer tokens across its own evaluations, with Kimi K3's reasoning shortened by 35% to 50% without sacrificing accuracy across seven benchmarks and two customers' production traffic. In one published production A/B run, reasoning tokens dropped 71.3% and total tokens dropped 39% at essentially unchanged task scores.

ShareShare on XShare on LinkedIn
← Previous storyH Releases Holo4 Open Models for GUI, Code, and Model Context Protocol

Related stories

  • Token & cost optimizationExtract LLM Logprobs to Run Single-Token Vision and Text Classification
  • Token & cost optimizationNVIDIA Cuts Confidential Computing Overhead in TensorRT-LLM Below Five Percent
  • Token & cost optimizationFast Jev Compaction Replaces Lossy Summaries in Claude Code
  • Token & cost optimizationAnthropic Open-Sources Claude-Generated Custom GPU Kernels for 4x Faster Inference

Email digest

Get the morning AI brief

One email a day — the stories that matter for engineers, founders and tech leads. Human-edited, with links to primary sources.

  • ✓120+ sources scanned daily
  • ✓Edited by a human
  • ✓1 email per day
  • ✓EN + UA

By subscribing you agree to the privacy policy.