ANALYSIS 11 min read

AI Science Agents Need an Evidence Ladder, Not a Discovery Score

Anthropic's ART result, Google's AI co-scientist, and published laboratory validations show a common pattern: agents can widen hypothesis search, but experiments determine what survives.

By EgoistAI ·
AI Science Agents Need an Evidence Ladder, Not a Discovery Score

AI-for-science announcements often compress several different achievements into one word: discovery. A model may retrieve a forgotten paper, generate a plausible hypothesis, rank candidates, predict an experiment, or help interpret a result. These are all useful. They are not equivalent.

Anthropic’s September 23 ART report provides a timely test case. Claude agents searched genomic data and flagged an unusual enzyme system; human scientists performed the laboratory work; the system’s function remains unknown. Google’s earlier AI co-scientist generated and ranked biomedical hypotheses, some of which collaborators tested in cells. Taken together, the projects suggest a clearer way to evaluate research agents: an evidence ladder.

Level one: retrieval and reproduction

The lowest useful level is not novelty. It is dependable reconstruction of what is already known. An agent should find the relevant literature, identify datasets and methods, reproduce established computational results, and show where its answer came from.

Anthropic describes this as a normal first step in its genome-mining workflow. Claude reads prior work and reproduces known families before looking for strange neighbors. That is more than a warm-up. Reproduction tests whether the system has understood database formats, sequence relationships, and the limits of the available evidence.

Failure here should block stronger claims. A system that cannot recover a known baseline is not ready to label an anomaly novel.

Level two: hypothesis generation

At the next level, an agent proposes explanations or candidates that are logically compatible with existing evidence. Large models are naturally prolific, which is both an advantage and a problem. Producing one thousand plausible hypotheses is easy compared with identifying the one experiment worth running.

Google’s AI co-scientist uses specialized generation, reflection, ranking, evolution, proximity, and meta-review agents. The system debates and iteratively improves proposals. Anthropic used hundreds of agents to filter more than 200,000 reverse transcriptases into thousands of candidates, then 20 reports.

The meaningful metric is not hypothesis count. It is the precision of the shortlist under expert review: how many candidates are genuinely novel, testable, safe, and valuable enough to justify scarce laboratory resources?

Level three: independent computational checks

A candidate becomes stronger when separate methods point in the same direction. That may include a literature search for prior art, alternative sequence-alignment tools, structure prediction, held-out data, sensitivity analysis, or an adversarial agent asked to disprove the idea.

Multi-agent debate can help, but model-on-model agreement is not independence. Two agents using the same base model and context can share the same blind spot. Strong systems vary tools, data slices, prompts, and evaluators, then preserve disagreement rather than averaging it away.

This is also where provenance matters. A research report should expose database versions, code, intermediate outputs, model versions, prompts, and the reason rejected candidates were discarded. Otherwise the final narrative can look cleaner than the search actually was.

Level four: real-world validation

The evidence changes category when a physical experiment tests a prediction. Google reported laboratory validation for selected drug-repurposing proposals in acute myeloid leukemia cell lines and experimental work on other biomedical questions. Anthropic reports that its team expressed and characterized elements of ART and observed distinct short RNAs from the repeat array.

An in-vitro result is still not a therapy, organism-level mechanism, or clinical outcome. It does, however, separate a computational story from a measurable intervention. The ladder should record the exact step reached: simulated, retrospective, in vitro, in vivo, prospectively replicated, or clinically validated.

Level five: independent replication and use

The strongest evidence arrives when an outside group reproduces the result and the method continues to work beyond the original showcase. Peer review can improve scrutiny, but publication alone is not replication. A deployed tool also needs prospective performance, failure reporting, and evidence that it improves decisions rather than merely generating more text.

Neither ART nor Google’s highlighted co-scientist validations should be treated as proof that a general agent can autonomously solve arbitrary research problems. They are evidence that particular workflows, teams, datasets, and experiments produced promising results.

Why the ladder changes product design

A system optimized for a single “discovery score” will learn to create confident novelty narratives. A system optimized for progress up an evidence ladder has different incentives. It must cite prior work, state what would falsify a proposal, choose discriminating experiments, track negative outcomes, and stop when evidence is insufficient.

This also clarifies human roles. Scientists are not merely approving the final answer. They define worthwhile questions, notice when a benchmark is a poor proxy, judge experimental feasibility, manage safety, and interpret ambiguous observations. The agent expands search and records reasoning; the team owns epistemic and physical consequences.

A practical evaluation protocol

For each agent-generated claim, record five fields:

  1. Evidence level: retrieval, hypothesis, computational corroboration, physical validation, or independent replication.
  2. Novelty check: databases and dates searched, plus related work found.
  3. Falsifier: the observation that would count against the claim.
  4. Human intervention: where experts changed goals, code, ranking, or interpretation.
  5. Cost per survivor: compute, review hours, laboratory time, and number of candidates rejected.

Publish the denominator. If one successful candidate emerged from 3,500 reports, that funnel is not an embarrassment; it is central evidence about the system’s precision and economics. Negative campaigns also matter because they reveal whether the agent knows when there is nothing worth testing.

Limitations

The available examples come largely from organizations building the models, and both teams selected projects compatible with computational search plus laboratory validation. Other sciences may have slower feedback, noisier measurements, inaccessible data, or ethical constraints that make the same architecture unsuitable.

Agent counts and token totals are not measures of scientific quality. More parallel search may improve recall while flooding reviewers with correlated errors. Automated rankings can become self-reinforcing when the generator and evaluator share training data or stylistic preferences.

Finally, a ladder does not eliminate judgment. Moving from a cell assay to an animal study or from correlation to mechanism requires domain-specific standards. The framework is a reporting discipline, not a universal formula.

Final verdict

Research agents are starting to produce outputs that survive contact with experiments. That is a meaningful threshold. The responsible way to describe progress is not to ask whether an AI “made a discovery,” but to show how far each claim climbed: from retrieval, to hypothesis, to corroboration, to experiment, to independent replication. The ladder makes both the achievement and the remaining uncertainty visible.

Share this article

> Want more like this?

Get the best AI insights delivered weekly.

By subscribing, you agree to our Privacy Policy. You can unsubscribe at any time.

> Related Articles

Tags

AI for scienceresearch agentsscientific methodevaluationlaboratory automation

> Stay in the loop

Weekly AI tools & insights.