08/23/2026 | Press release | Distributed by Public on 08/23/2026 18:18
Inherent, a London-based artificial intelligence startup founded by former Google DeepMind researchers, says its new AI agent has outperformed much larger systems from OpenAI and Anthropic in a test designed to measure whether AI can independently reproduce scientific research.
The startup, which emerged from stealth just weeks ago with a $50 million seed funding round, said its agent Faraday surpassed Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5 in a benchmark focused on reproducing the findings of published scientific papers without being given the expected results in advance.
The result is notable not simply because Faraday outperformed two larger frontier models, but because Inherent said the agent runs on Qwen 3.6, a comparatively small model with 27 billion parameters.
Register for the next Tekedia Mini-MBA.
Register for Tekedia AI in Business Masterclass.
Join Tekedia Capital Syndicate and co-invest in great global startups.
Register for Nigeria Capital Market Masterclass.
Parameters are a broad measure of the number of learned values in an AI model and are often associated with model size, although parameter count alone does not determine a system's capabilities, training cost, or efficiency.
For Inherent, the more important achievement is how Faraday reaches its conclusions. The company is pursuing a much broader objective than simply reproducing existing scientific findings. Its long-term ambition is to develop AI agents capable of discovering new scientific knowledge and contributing to research across multiple disciplines.
Edward Hughes, Inherent's cofounder and chief scientist, said reproducing published research is an important starting point because it is also a common exercise for human researchers.
"Many PhD students actually start by doing this," Hughes said.
The company therefore views paper replication as a test of whether an AI system can independently formulate experiments, execute them and interpret the results rather than simply answer questions based on information already contained in its training data.
"What was most interesting to us about this was not so much the result of beating those frontier agents - which of course we liked - but was actually the way we went about building this," Hughes told TechCrunch.
Inherent said it also set a higher bar than simply measuring whether Faraday could reproduce published results. The company wanted the agent to demonstrate what it calls "research taste," meaning an ability to identify worthwhile questions, determine which experiments are useful, and design those experiments effectively.
That capability is difficult to encode through conventional instructions because it involves judgment about which research directions are likely to produce useful information.
Inherent uses reinforcement learning to address that problem. Instead of attempting to explicitly teach the agent every step involved in scientific research, the company rewards the system for producing desirable outcomes and allows it to learn strategies that lead to those outcomes.
The approach is central to Inherent's broader thesis that an AI scientist should develop transferable research capabilities rather than simply memorize procedures for particular scientific fields.
"We're always guided by that north star of building an AI scientist agent and imbuing our agents with taste," Hughes said.
That philosophy has also influenced what Inherent has chosen not to build. Rather than developing its own coding system, Faraday uses OpenAI's GPT-5.5 Codex for software development tasks. The company compares that approach with how human scientists work, relying on existing tools rather than attempting to build every piece of software needed for an experiment.
The strategy could make a huge difference as AI research systems become more specialized. Instead of competing with every major AI developer on the underlying model, Inherent is attempting to build an agentic layer capable of combining models and tools to perform complex scientific work.
Hughes said the company also wants Faraday to behave more like a research collaborator than an AI assistant designed primarily to satisfy its user. The goal, he said, is an agent that can independently investigate a question and return with unexpected findings rather than simply confirming what the user already believes.
That is expected to become more useful as AI systems move from generating answers to carrying out autonomous research. A useful scientific agent needs to be capable of challenging assumptions, pursuing alternative hypotheses, and reporting results that may contradict the user's expectations.
Inherent's operating model is similarly focused on maintaining a small, concentrated research team. Its roughly dozen employees currently work in person from an office in London's King's Cross, an area that has developed into a major AI research and startup hub partly through the presence of Google DeepMind.
"We believe that London is the place to be," Hughes said.
The company is nevertheless critical of one aspect of Britain's employment system that can make it harder for startups to recruit experienced AI researchers.
Hughes has called for an end to "garden leave," a practice under which employees can be prevented from joining a competitor or starting a competing company for a period after leaving their previous employer.
He said that the practice can put British AI startups at a disadvantage compared with companies in the United States, where researchers generally face fewer restrictions when moving between employers.
"This is a personal view rather than a company view, but I was affected by the garden leave problem," Hughes said.
Hughes eventually overcame the restriction and founded Inherent with two other former DeepMind employees and a fourth cofounder.
The startup now plans to increase its workforce to between 20 and 25 employees by the end of the year. Its ambitions extend beyond scientific agents into world models, potentially putting it in competition for talent with much larger AI laboratories.
That hiring push could become a major boost as researchers reassess their positions at established AI labs. Demis Hassabis, DeepMind's cofounder and CEO, has taken on a new role, while changes across the broader AI industry are creating opportunities for researchers to move into startups.
Inherent's early benchmark results do not establish that a 27 billion-parameter model is generally more capable than much larger frontier systems. The test covers a specific scientific-research task, and performance on paper replication does not necessarily translate into broader reasoning, coding, or general-purpose capabilities.
But the result underpins that in AI development, raw model size may not be the only route to stronger performance on complex tasks. Inherent is betting that reinforcement learning, tool use, autonomous experimentation, and specialized agent architecture can allow relatively small underlying models to perform sophisticated research tasks.
If that approach generalizes beyond reproducing existing scientific work, the implications could be significant. Instead of simply making AI models larger, developers may focus more on teaching smaller systems how to choose problems, conduct experiments, use external tools, and learn from the results.