London startup's small AI agent outperforms larger OpenAI and Anthropic systems in science test

London AI startup Inherent says its 27-billion-parameter agent Faraday beat OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.8 at reproducing published scientific research. The company, backed by $50 million in seed funding, argues agent design and training can matter more than raw model size.

Categorized in: AI News Science and Research
Published on: Aug 24, 2026
London startup's small AI agent outperforms larger OpenAI and Anthropic systems in science test

London-based AI startup Inherent says its newly released agent, Faraday, outperformed larger systems from OpenAI and Anthropic in reproducing published scientific research - despite running on a model with just 27 billion parameters. The company, founded by former Google DeepMind researchers, is developing AI systems meant not simply to answer questions but eventually to contribute to scientific discovery. Its early results suggest that agent design and training techniques can sometimes matter more than raw model size.

Inherent said Faraday outperformed Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5 in independently reproducing findings from published scientific papers. The agent was not given final answers in advance. It had to interpret the research, design experiments, and attempt to reproduce the reported findings.

Scientific replication is a standard part of research training. Inherent co-founder and chief scientist Edward Hughes said many PhD students begin their research careers by reproducing previous experiments before moving on to original work. For Inherent, however, replication is an intermediate step toward a larger goal: building an AI scientist capable of discovering new knowledge.

A 27-billion-parameter model competes with frontier systems

One of the most striking aspects of Faraday's performance is its underlying model. While GPT-5.5 and Claude Opus 4.8 are large frontier systems, Faraday runs on Qwen 3.6 with approximately 27 billion parameters. Parameter count is commonly used as a rough indicator of model size and the computing resources required for training and inference.

That raises a question for the AI industry: could better training and agent architecture compensate for a smaller underlying model? Hughes said beating larger frontier agents was not the most important part of the experiment. The company was more interested in the approach it used to build the system.

Inherent wants AI to develop "research taste"

Accuracy alone was not enough for Inherent. The company also wanted Faraday to demonstrate what it calls "research taste" - the ability to determine which experiments are worth conducting, which questions are important, and how scientific investigations should be designed.

That type of judgment is difficult to define using fixed rules. Instead, Inherent relies heavily on reinforcement learning, a training technique in which an AI system receives rewards for achieving desirable outcomes rather than being explicitly instructed how to behave in every situation. The company believes this approach could help its agents develop skills that generalize across scientific disciplines.

Rather than primarily training Faraday on descriptions of how scientists work, Inherent focuses on rewarding successful research behavior. The long-term objective is to create agents capable of deciding what to investigate, running experiments, analyzing results, and identifying promising directions. Hughes describes the goal as building an AI scientist that behaves more like a capable research collaborator than a traditional chatbot - one that might independently investigate an unexpected result and return with new experiments and observations.

Inherent is also avoiding the temptation to build every component of its stack internally. Rather than developing its own coding model, the company allows Faraday to use OpenAI's GPT-5.5 Codex for programming tasks. The approach mirrors how human scientists work: researchers typically rely on existing software and tools rather than building everything from scratch. For Inherent, the value lies in developing an agent capable of selecting and using the right tools effectively.

From assistants to autonomous scientists

Most AI assistants today are designed to respond directly to user instructions. Inherent is pursuing something more autonomous: agents that investigate ideas independently, design useful experiments, and challenge assumptions rather than simply providing answers users expect to hear.

That distinction could become increasingly important as AI agents move from administrative and coding tasks into scientific research. An effective AI scientist would need not only strong reasoning capabilities but also curiosity, experimentation, judgment, and the ability to identify important questions.

Faraday's release comes shortly after Inherent emerged from stealth with a $50 million seed funding round. The startup was founded by former Google DeepMind researchers and currently has around a dozen employees working from its office in King's Cross, London. The area has become one of Europe's most important AI hubs, supported in part by the presence of Google DeepMind and a growing concentration of AI companies. Inherent plans to expand to roughly 20 to 25 employees by the end of the year.

Hughes believes London's concentration of AI researchers gives companies like Inherent a strong advantage. The startup could also become an attractive destination for researchers leaving larger AI laboratories, particularly as competition for specialized AI talent intensifies. However, Hughes has criticized the U.K.'s practice of "garden leave," which can prevent employees from immediately joining or launching competing companies after leaving their employer. He argues that such restrictions may make it harder for British startups to recruit experienced researchers as quickly as their U.S. competitors.

Beyond AI agents for scientific research, Inherent has ambitions in world models - systems that aim to help AI build richer internal representations of environments, relationships, and cause-and-effect dynamics. Combining these capabilities with autonomous research agents could eventually enable AI systems to reason about scientific problems in more sophisticated ways. The broader vision is an AI system that does not merely retrieve existing knowledge but can generate hypotheses, test them, and potentially contribute new discoveries.

Faraday's results add to a growing debate about whether AI progress will always depend on building increasingly large models. If specialized agents running on smaller models can outperform frontier systems in specific tasks, the future of AI competition may depend less on raw parameter counts and more on training methods, tool use, reasoning strategies, and agent architecture. For startups, that could be particularly significant: building frontier-scale foundation models requires enormous capital and computing infrastructure, while more efficient approaches could allow smaller companies to compete by focusing on specialized capabilities.

For professionals working in science and research, the practical takeaway is that model size is no longer the only meaningful benchmark. Faraday is still focused on replicating existing work rather than making independent discoveries, but it demonstrates that a well-trained, smaller agent can handle complex research tasks - and that skills like experimental design and judgment may matter more than raw compute. Those evaluating AI tools for research should pay attention to agent architecture and training methods, not just parameter counts. Inherent's broader AI for Science & Research ambitions suggest the gap between AI assistants and autonomous research collaborators may close faster than many expect, and AI Research Courses may soon need to cover agent design alongside model capabilities.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)