Science AI
FirstResearch Framework Makes AI-Generated Scientific Questions Auditable
A new framework called FirstResearch introduces a structured Research Question Certificate to make the first research question proposed by LLM scientific discovery agents inspectable before downstream execution.
The certificate records primitive definitions, assumptions, a mechanism model, a tension or contradiction, a falsifiable hypothesis, a minimal decisive test, and a failure update rule. On ten LLM-agent research topics, FirstResearch outperforms controlled prompt-level baselines inspired by AI co-scientist, Agent Laboratory, and AI Scientist-v2 under a primary DeepSeek-blind-judge protocol.
A Gemini-2.5-Flash independent-judge rescore of the same 40 baseline packages preserves the system-level ranking, with FirstResearch scoring 4.86/5 versus 4.38/5 for the strongest baseline and Pearson agreement of 0.865 on average score. A one-repeat ablation checkpoint suggests the certificate-centered core is the strongest component: certificate-only scoring reaches 4.90/5 under DeepSeek and 4.88/5 under Gemini, while removing certificates drops below 1/5 under both judges.
Results are preliminary and use LLM judges rather than human domain experts, but support the claim that explicit derivation constraints are a promising mechanism for making LLM-generated scientific questions more auditable.
Sources
Evidence entered
Admission Evidence and chronology passed Science AI
Publication receipt Entered the validated Newswire
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire

