Research agents should make hypotheses easier to challenge.

What Co-Scientist and ChemCrow suggest about tool-assisted research, and a proposed evidence workflow that preserves uncertainty and expert review.

A source-linked editorial perspective. Research findings are attributed below; deployment blueprints and evaluation plans are Stellitron proposals, not customer case studies or promised outcomes.

Generation and validation are separate tasks

PUBLISHED RESEARCH / DOCUMENTATION

The Co-Scientist paper, first submitted in 2025 and revised in June 2026, describes a multi-agent system for generating, critiquing and refining research hypotheses. Its validation focuses on selected biomedical applications. The findings concern the authors’ system and experiments; they do not prove general discovery ability or clinical effectiveness for a new deployment.[1]

Specialist tools change the research workflow

PUBLISHED RESEARCH / DOCUMENTATION

ChemCrow studies an LLM agent connected to expert-designed chemistry tools. The authors also report limitations in model-based evaluation, including an evaluator’s difficulty distinguishing wrong model outputs. This supports careful investigation of tool-assisted research and independent evaluation, rather than treating an eloquent generated explanation as scientific evidence.[2]

Build an evidence ledger first

PROPOSED BLUEPRINT

A first project could support a materials team reviewing a narrowly defined question. Gather permitted primary papers, attach stable citations and record the experimental setting, methods, observations and limitations. Preserve contrary findings instead of asking a summariser to resolve every disagreement.

Maintain separate fields for published findings, an analyst’s interpretation and a proposed hypothesis. A missing measurement should remain missing. Keep source versions and retrieval dates so an expert can revisit the evidence when a paper changes or a conclusion is disputed.

Make critique an inspectable step

PROPOSED BLUEPRINT

Use one stage to produce an evidence summary and another to search for contradictory observations or missing conditions. Ask the reviewing researcher to judge whether each proposed hypothesis follows from the cited material and which experiment would distinguish competing explanations. Agreement between agents is not independent experimental confirmation.

The deliverable is a review packet with citations, uncertainty, alternatives and a proposed investigation. Any laboratory execution, biological intervention or clinical use is outside this initial workflow and requires domain-specific professional oversight. This blueprint is hypothetical; it claims no Stellitron research discovery.

Use experts and held-out questions

PILOT EVALUATION

Evaluate source fidelity, citation validity, representation of contrary evidence and the usefulness of suggested next questions. Have subject-matter experts review blinded outputs alongside the existing literature-review process. Include misleading abstracts, withdrawn claims and papers whose methods do not match the target setting.

Record the time to check the evidence as well as the time to generate it. An agent that rapidly creates hypotheses but makes verification harder may not improve the workflow. Choose further automation based on measured research utility, with scientific validation remaining distinct from generated content.

Sources & publication dates

Primary papers and official documentation consulted for this perspective. Source publication dates differ from the date of this article; undated documentation is identified explicitly.

  1. Accelerating scientific discovery with Co-Scientist (v2)Source published 29 June 2026 · Consulted 3 October 2026
  2. ChemCrow: Augmenting large-language models with chemistry toolsSource published 11 April 2023 · Consulted 3 October 2026

Test a first direction.

Bring your own constraints into the demo. Explore an initial workflow, then discuss the evidence, integrations and evaluation a pilot would need.

Explore this workflow

Browse architecture concepts and decks

Research agents should make hypotheses easier to challenge. | Stellitron