跳到主要內容
AI News HubLIVE
站內改寫3 分鐘閱讀

待翻譯:Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Fireworks AI has released Ember-1, a post-trained Kimi K3 that learns to produce shorter reasoning traces instead of lowering reasoning effort. Fireworks reports about 40% fewer tokens, with output tokens per task falling from 49.3K to 29.9K in a production A/B test at an essentially unchanged score. Ember-1 is available now as an API-only Research Preview at Kimi K3 pricing. The post Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens appeared first on MarkTechPost.

來源MarkTechPost作者: Asif Razzaq
待翻譯:Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens
回報錯誤

更正管道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

Fireworks AI has released Ember-1, a specialized model from Fireworks Research built by post-training Moonshot AI’s open-weight Kimi K3. Ember-1 learns to produce shorter reasoning traces while keeping task accuracy. This is different from lowering the reasoning effort setting at inference time. According to the Fireworks release post, Ember-1 delivers Kimi K3’s quality with about 40% fewer tokens. Is it deployable? Yes, but only through the Fireworks serverless API as a Research Preview. Fireworks has not released Ember-1’s weights, training code, or exact training algorithms, so self-hosting is not an option today. The Problem: Reasoning Models Think Too Much Fireworks team reports that reasoning models like Kimi K3 sometimes spend more than 90% of generated tokens on internal reasoning. That cost compounds in multi-turn agentic workloads. Each turn replays prior reasoning back to the model. Context grows roughly quadratically with the number of turns. Long traces from early turns get re-read, and re-billed, on every later call. Fireworks team explains how customers wanted K3’s coding capability at lower cost. Turning down K3’s reasoning effort did not solve it. Lower effort settings gave up too much quality. So the team trained the model to reason more efficiently instead. How Fireworks Research Built Ember-1 Not all of K3’s reasoning is waste. Some of it is useful self-reflection, like revisiting an assumption or reacting to feedback. Ember-1 keep that behavior while cutting redundant reasoning and unproductive loops. The training collection spans mathematics, coding, instruction following, conversation, search, tool use, and software engineering. It covers both standalone problems and extended multi-step interactions. Task and environment feedback guides on-policy planning and learning. Fireworks team ran more than 50 training experiments and over 200 evaluations. They also developed new training algorithms, which it has not published. All training ran on Fireworks Serverless Training. Fireworks states it used its own data and no customer data. Benchmark Results Fireworks compared Ember-1 with Kimi K3 at three reasoning effort levels. Cost was computed with public Kimi K3 API pricing. These are Fireworks’ own published evaluations. BenchmarkNK3 LowK3 HighK3 MaxEmber-1Ember-1 vs K3 Max (cost) Terminal Bench 2.18976.4%77.6%80.9%82.0%-51.9% / -23.1 USD SWE-bench Verified50080.4%86.0%93.2%92.2%-15.5% / -68.1 USD SWE-Interact756.7%13.3%21.3%20.0%-32.5% / -60.8 USD DeepSWE 1.111355.8%62.8%66.4%75.2%-23.7% / -126.9 USD τ-2 Bench Airline5064%64%64%66%-5.9% / -0.3 USD Ember-1 leads K3 Max on Terminal Bench 2.1 and DeepSWE 1.1. It trails slightly on SWE-bench Verified and SWE-Interact. Across seven benchmarks and two customers’ production traffic, Fireworks says K3’s reasoning was shortened by 35 to 50% without sacrificing accuracy. On Doximity’s Bedside Bench, a physician-validated set of 500 clinical cases, Ember-1 set a new cost-per-task Pareto frontier. That result comes from Fireworks’ new Specialized Intelligence Index. Production A/B Test Results Fireworks ran live A/B tests with 2 customers on production coding workloads. Both saw roughly 35% fewer tokens per task at comparable quality. In the published run, output tokens fell from 49.3K to 29.9K per task. Reasoning tokens dropped 71.3% and total tokens dropped 39%. The task score was essentially unchanged: 0.753 for Ember-1 versus 0.751 for K3. Average steps fell from 23.8 to 21.4. One customer now runs Ember-1 in production. Ember-1 costs the same per token as Kimi K3 on Fireworks: $3.00 input, $0.30 cached input, and $15.00 output per 1M tokens. The savings come entirely from generating fewer tokens. At that output rate, the A/B figures work out to about $0.74 versus $0.45 in output cost per task (our calculation, output only). Interactive Explainer Key Takeaways Ember-1 is Kimi K3 post-trained to reason in fewer tokens, not run at lower effort. Fireworks reports about 40% fewer tokens with accuracy held across its evaluations. In a production A/B test, output fell from 49.3K to 29.9K tokens per task at a 0.753 vs 0.751 score. It beats K3 Max on Terminal Bench 2.1 (82.0%) and DeepSWE 1.1 (75.2%), trailing on SWE-bench Verified (92.2%). API-only Research Preview at K3 pricing; weights and training code are not released. Check out the Technical Details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens appeared first on MarkTechPost.

展開要點與分析

文章情報

工程師進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級後設資料。
  • Fireworks AI has released Ember-1, a post-trained Kimi K3 that learns to produce shorter reasoning traces instead of lowering reasoning effort. Fireworks reports about 40% fewer t…

技術影響

可能影響 Agent 架構、工具呼叫、工作流自動化和產品整合。

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。