
Inherent’s small AI agent beats frontier models at paper replication

Inherent, a London AI lab founded by Google DeepMind alumni, says its AI agent Faraday outperformed Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 at independently reproducing the findings of published scientific papers without being given the answer in advance. The company claims this result was achieved with a far smaller model: Faraday runs on Qwen 3.6, a 27-billion-parameter model, compared to the much larger frontier systems. Chief scientist Edward Hughes emphasized that beating larger agents was not the primary goal; the more interesting part was the training approach.
Rather than training agents primarily on examples of how science is conducted, Inherent uses reinforcement learning to reward the agent for good experimental outcomes. The company’s stated aim is to build an ‘AI scientist agent’ that can discover new knowledge, not just verify existing results. To that end, Inherent set a higher bar than accuracy alone: Faraday must demonstrate ‘research taste’ — an instinct for which experiments are worth running and how to design them. Hughes said this taste is what reinforcement learning is meant to instill, and the team believes this approach will generalize better to their longer-term goal of contributing across multiple scientific fields.
The design choices reflect that north star. Inherent avoids building its own coding tool and instead had Faraday use OpenAI’s GPT-5.5 Codex, much as human scientists rely on existing software. The company also tries to avoid building agents that simply tell users what they want to hear; the ideal teammate, Hughes said, is one who comes back with curiosity-driven experiments and asks for feedback.
Inherent emerged from stealth just weeks ago with a $50 million seed round. It has about a dozen employees, all working in person in an office in King’s Cross, a London neighborhood where Google DeepMind’s presence has helped create a major AI hub. Hughes is bullish on London’s AI talent density, but he also spoke out against ‘garden leave’ — the U.K. practice of barring departing employees from joining or starting a rival company for months after resignation. He said he was personally affected by this constraint before he and three co-founders launched Inherent. The startup plans to grow to about 20 to 25 employees by year-end and is also pursuing world models, which could make it an attractive destination for DeepMind staff unsettled by Demis Hassabis’s new role.
The reported performance claim rests on Inherent‘s own evaluation, and the company has not released independent benchmarks. Hughes acknowledged that the result of beating frontier agents was welcome but secondary; the significance, he argued, lies in the training methodology and its potential to scale toward genuine scientific discovery. The source provides no quantitative details on the replication task, so the practical margin of improvement over Claude Opus 4.8 or GPT-5.5 is not documented here. Inherent‘s approach deliberately prioritizes reward-based generalization over explicit instruction, and it is selective about which tools to build internally, a stance that will be tested as the lab scales its headcount and research ambitions.


