LIVE
+++ Nvidia agent AVO solves all 183 ARC-AGI-3 levels — 100 percent +++ DeepSeek launches vision model V4-Flash-Vision-Exp at a cut price +++ Anthropic hires Google TPU architect Amir Salek for in-house chips +++ Broadcom seeks over $60 billion in debt for Anthropic chips +++ Reversal: OpenAI wants California's AI law SB 53 tightened +++ London startup Inherent: agent Faraday beats frontier models at research replication ++++++ Nvidia agent AVO solves all 183 ARC-AGI-3 levels — 100 percent +++ DeepSeek launches vision model V4-Flash-Vision-Exp at a cut price +++ Anthropic hires Google TPU architect Amir Salek for in-house chips +++ Broadcom seeks over $60 billion in debt for Anthropic chips +++ Reversal: OpenAI wants California's AI law SB 53 tightened +++ London startup Inherent: agent Faraday beats frontier models at research replication +++
All news ›
AI IN LIFE AI IN LIFENEWS
DAILY
RESEARCH

Faraday: small AI agent beats frontier models in the lab

London startup Inherent — founded by DeepMind alumni — says its agent Faraday beat Opus 4.8 and GPT-5.5 at replicating published research.

Faraday: small AI agent beats frontier models in the lab

Illustration · AI-generated (AI IN LIFE)

At a glance

  • Inherent claim of August 22, 2026: Faraday beats Opus 4.8 and GPT-5.5 at research replication — with no published scores
  • Base: Qwen 3.6 with 27 billion parameters plus GPT-5.5 Codex for coding
  • Training primarily via reinforcement learning, focused on "research taste"
  • Founded by DeepMind alumni; $50 million seed, stealth exit May 2026
  • Team: about 12 employees, planned 20–25 by year-end

What is claimed? London startup Inherent says its AI agent Faraday outperformed both Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5 at replicating published research papers — reported by TechCrunch on August 22, 2026. The company has not published concrete benchmark scores, however — the claim cannot currently be verified independently.

Who is behind it? Inherent was founded by Google DeepMind alumni Edward Hughes (chief scientist), Louis Kirsch, Kaloyan Aleksiev and Tantum Collins, is based in King's Cross, and emerged from stealth in May 2026 with a $50 million seed round. The team counts about 12 people and plans to grow to 20–25 by year-end.

What makes Faraday technically notable? The agent runs on Qwen 3.6 with 27 billion parameters — tiny compared with the frontier systems it reportedly beat — and uses GPT-5.5 Codex for coding tasks. Training relied primarily on reinforcement learning: Faraday is meant to develop "research taste", the instinct for which experiments are worth running and how to set them up properly.

How does the founder frame it? Hughes himself plays down the benchmark win: "What was most interesting to us about this was not so much the result of beating those frontier agents — which of course we liked — but was actually the way we went about building this."

Why does it matter? If confirmed, the result would be further evidence that specialised small agents can beat large generalists in niches — at far lower compute cost. For research departments and science-adjacent companies, a new tool class emerges: the AI colleague that reproduces studies before you build on them.

◈ AI-GENERATED REPORT · SOURCES LINKED

FAQ

Are the results independently confirmed?

No. Inherent has not published concrete benchmark scores; the claim rests on company statements.

What is research replication?

An agent tries to independently reproduce the experiments and results of a published paper — a hard test of scientific work.

Why is the model size remarkable?

Faraday uses a 27-billion-parameter model yet reportedly beat much larger frontier systems — a signal about efficiency.