Fireworks releases Ember-1 reasoning model on AI Gateway

Fireworks AI released Ember-1, a reasoning model that generates roughly 40% fewer tokens than its Kimi K3 base at comparable quality. It's available through Vercel's AI Gateway with a 1M-token context window and a two-week preview.

Published on: Sep 28, 2026
Fireworks releases Ember-1 reasoning model on AI Gateway

Fireworks AI has released Ember-1, a research preview reasoning model built on Kimi K3, now available through Vercel's AI Gateway. The model targets coding and agentic workflows, producing roughly 40% fewer generated tokens than its base model at comparable quality - a difference that can lower output costs and reduce context bloat in multi-step agent tasks.

The model launches with a two-week preview window and arrives with a 1M-token context window, support for text and image inputs, tool calling, and implicit prompt caching. Fireworks' endpoint enforces Zero Data Retention and No Prompt Training policies.

What the model offers for developers and agents

Ember-1's primary selling point is efficiency in reasoning traces. For coding agents that make repeated model calls, shorter outputs mean less accumulated context and lower per-call costs. Fireworks reports that the quality remains comparable to Kimi K3 across its internal evaluations while generating fewer tokens.

Developers can access the model through the standard AI SDK by specifying fireworks/ember-1 as the model name. The setup path through Vercel CLI is straightforward: install the latest CLI, run vercel ai-gateway setup, and the command detects installed coding agents, provisions an API key, and configures the connection. From there, selecting the model in the agent's configuration is all that's required.

AI Gateway integration

AI Gateway provides a unified API layer for calling models across providers, with baked-in usage tracking, cost monitoring, API key budgets, and routing rules. Making Ember-1 available through this gateway means teams can test the model alongside their existing stack without changing authentication flows or monitoring tools.

The model can also be tested directly in the model playground, which offers a low-friction way to evaluate output quality before committing to integration work.

Why this matters for finance, IT, and product teams

For teams building internal tools or customer-facing agents, token efficiency translates directly to operating costs. A 40% reduction in generated tokens per call, maintained across dozens or hundreds of daily agent invocations, compounds quickly. In finance and product development contexts where LLM costs are line items in project budgets, models that deliver comparable quality with fewer tokens deserve evaluation.

IT and development leads evaluating agent frameworks should note the implicit prompt caching and tool-calling support. These features reduce the engineering overhead required to build reliable multi-step workflows. The two-week preview window is short, so teams interested in benchmarking should move quickly. Those looking to build stronger foundations in agent development can explore AI for Software Developers Courses or browse AI Agent Courses for practical implementation patterns.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)