HyperAIHyperAI

Command Palette

Search for a command to run...

LLM
Generative AI

Fireworks Unveils Ember-1 AI Model That Cuts Tokens By Half

Fireworks Research has introduced Ember-1, a specialized artificial intelligence model built on the Kimi K3 architecture that reduces token consumption by approximately 40 percent while maintaining answer quality. This release addresses the high costs associated with long reasoning traces in reasoning models, which can consume over 90 percent of tokens and become exponentially expensive in multi-turn agentic workloads. Ember-1 learns to eliminate unnecessary computation while preserving critical self-reflection, enabling efficient reasoning across diverse tasks. The development process involved over 50 training experiments and 200 evaluations, resulting in novel algorithms for token efficiency. Training encompassed software engineering, mathematics, instruction following, tool use, and extended interactions. Fireworks utilized its Serverless Training infrastructure to accelerate research and deployment. Performance assessments confirm Ember-1's superior cost-quality trade-off. On the Specialized Intelligence Index, Ember-1 set a new Pareto frontier for cost versus performance on Doximity's Bedside Bench, outperforming models including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 on cost-per-task metrics. In software engineering evaluations such as SWE-bench Verified, Ember-1 matched Kimi K3's maximum reasoning accuracy at a fraction of the cost, strictly dominating lower-effort K3 configurations. Token reductions ranged from 35 to 50 percent across benchmarks, with cost savings reaching 68 dollars per task on SWE-bench Verified. Live validation supports these results. Customer A/B tests on production coding workloads demonstrated a 35 percent reduction in token usage per task, with Ember-1 sustaining or improving task completion and success rates. One customer has moved Ember-1 to live production with plans to replace base models. Internal testing by Fireworks developers also confirmed seamless operation with no quality degradation alongside significant token savings. Ember-1 is available now as a Research Preview on Fireworks Serverless. The launch initiates a series of specialized models focused on efficiency. Fireworks is also releasing training support, allowing enterprises to build customized, token-efficient models tailored to specific workloads using proprietary data, reinforcing the trend toward specialized intelligence in the open model ecosystem.

Related Links