Subconscious raises $5.1 million to reduce the cost of running long AI agents

Subconscious raised $5.1M for an inference platform that claims to cut token costs by up to 80%. One customer reported monthly AI spending dropping from $40,000 to $6,000 after switching.

Published on: Sep 26, 2026
Subconscious raises $5.1 million to reduce the cost of running long AI agents

Subconscious, a Cambridge startup building an inference platform for long-running AI agents, has raised $5.1 million in combined pre-seed and seed funding. The company says its runtime can cut token costs by up to 80%, targeting engineering teams that want to run open-model agents more cheaply without swapping their existing tools or hardware.

Co-founders Jack O'Brien and Hongyin Luo announced the funding on September 22nd, 2026. The service is available to engineering teams through cloud and on-premises deployments. O'Brien said the company had been operating for more than a year and a half before the managed-service launch and funding announcement.

MIT research meets the cost problem

Subconscious grew out of inference optimization research at MIT. Luo holds a PhD from MIT's electrical engineering and computer science department and co-authored a 2025 paper with O'Brien on extending reasoning beyond conventional context limits. O'Brien's thesis is straightforward: open-weight models have become good enough for coding work, but the cost and speed of running them across long agent sessions remain barriers to adoption.

MassVentures led both funding rounds. Other participants include Foothill Ventures, Underscore VC, E14 Fund, Oakseed Ventures, Agent Fund, Companyon Ventures, and Taihill Venture. The company did not disclose a valuation or a breakdown of the pre-seed and seed rounds.

Compression and caching as the pitch

The platform combines dynamic context compression with caching to reduce the amount of context an agent must repeatedly process. This matters because software agents make many model calls over extended tasks, accumulating large token volumes. Subconscious reports that for workloads above 200,000 tokens, its platform can make tasks twice as fast, expand effective context beyond five million tokens, and cut costs by up to 80%. These are company-reported claims, not independently verified results.

Subconscious offers its own benchmark data. On the TriE systems benchmark, the company says its runtime completed tasks twice as fast as SGLang and handled 2.3 times as many concurrent requests. On DeepSWE, a long-coding-task benchmark, GLM 5.2 hosted on Subconscious solved 46% of problems at an average cost of $2.79, compared with 44% and $3.92 for the same model on what the company calls standard inference infrastructure. The announcement does not describe the full test setup, so these results should be read as vendor-produced claims.

The strongest commercial evidence comes from an unnamed customer account provided by the company. Subconscious says a 20-person engineering team switched from Claude to GLM 5.2 hosted on its platform in July, reducing monthly AI spending from $40,000 to $6,000 over the next two months. One engineer's agent trace ran for 4,571 turns and made 9,556 tool calls. Subconscious says its runtime recorded 449 million tokens, compared with 2.6 billion tokens that a conventional runtime would have billed. The company attributes the reduction to context compression and says the customer reported no loss in model capability.

An infrastructure play, not a new agent

Subconscious is selling an infrastructure layer that works across existing agent products. The announcement names Claude Code, Codex, Pi, Copilot, and OpenCode as tools that can connect through its CLI. The company is not asking customers to adopt a new agent interface. For teams that want to keep data and compute in their own environment, the service supports deployment on a customer's own GPUs.

O'Brien argues that open models' improving quality makes the timing right, especially as engineering teams impose spending limits even while developers use agents more heavily. The bet is that compression can make open-model agents cheaper to run without replacing customers' tools or hardware.

Why this matters for product and engineering teams

For teams already spending heavily on coding agents, Subconscious's cost claims - if they generalize - represent a direct path to lower inference bills without retooling. The unnamed team's reported drop from $40,000 to $6,000 per month is a specific data point, though it comes from a single customer and lacks independent verification. Engineers evaluating the platform will want to test whether compression delivers comparable savings on their own workloads and whether the claimed "no loss in model capability" holds up under their typical task complexity. The on-premises deployment option also gives teams a way to test this while keeping data inside their own infrastructure.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)