Smallest.ai has raised $13 million in Series A funding led by Seligman Ventures, with participation from Sierra Ventures and 3one4 Capital, bringing its total funding to over $21 million. The company also launched Hydra, an asynchronous speech-to-speech model built on its new Voice 4.0 architecture, designed to make voice AI interactions more natural and responsive by processing listening, reasoning, and speaking in parallel.
The global voice AI market is currently valued at $2.4 billion and is projected to reach $47.5 billion by 2034, according to Market.US. Yet AI accounts for less than 1% of all voice interactions today. Most deployments remain stuck in narrow customer service use cases because existing systems still sound robotic and suffer from noticeable latency.
The root cause is architectural. Voice AI today relies on a chained stack of separate components - speech recognition, large language model processing, orchestration layers, memory systems, text-to-speech engines, and guardrails - all cobbled together to execute sequentially. This disjointed pipeline introduces delays and makes interactions feel artificial.
A shift to asynchronous processing
Hydra, the foundational model inside Voice 4.0, breaks that chain. Instead of processing one step after another, the model handles listening, reasoning, taking actions, and responding in parallel. That allows real-time conversational flow, natural interruptions, and mid-conversation tool use. It can also pair with the company's earlier speech-to-text models for transcription with latency measured in milliseconds.
"Humans don't wait for someone to finish speaking before they begin thinking. We listen, think, and respond simultaneously," said Sudarshan Kamath, founder and CEO of Smallest.ai. "Voice AI needs to work the same way. By rethinking the stack instead of simply scaling models, we're reducing latency to the point where voice interactions feel genuinely human."
Existing traction and developer tools
Smallest.ai already offers speech-to-text models like Pulse STT Pro and Lightning V3.1, which rank among the top systems on the Artificial Analysis benchmark. When Lightning launched last year, it was described as the fastest text-to-speech model on the market, generating 10 seconds of speech in 100 milliseconds. The model now supports 38 languages and includes emotion detection, speaker diarization, data redaction, and noise reduction.
Customers such as RingCentral, Truecaller, Kogtal Financial, and Readymode use these tools to reduce customer support costs by up to 80% in some cases.
Competitive landscape
Smallest.ai is not alone in the race to improve voice AI. Earlier this week, Fish Audio raised $52 million for its open-source speech model platform. ElevenLabs, the best-funded player in the category, closed a $500 million round in February to build its agentic voice AI platform. Despite the competition, Seligman Ventures' Ashish Kakran sees room for multiple winners, especially if they simplify the developer experience.
"Developers now increasingly talk to their machines instead of typing code," Kakran said. "Smallest.ai is taking a fundamentally different approach to the category by rethinking architecture itself. Customers get an efficient vertically integrated stack and don't need to waste time stitching models together."
Why this matters for IT and development
For developers and IT teams, the move from sequential to asynchronous voice processing reduces the heavy lifting of integrating disparate models. A vertically integrated stack that handles speech recognition, reasoning, and synthesis in parallel means fewer points of failure, lower latency, and less custom orchestration code. This opens the door to voice-enabled applications that can respond mid-sentence, use tools on the fly, and handle real-world conversational dynamics - something chained architectures have struggled to deliver.
Your membership also unlocks: