# FirstResearch Framework Makes AI-Generated Scientific Questions Auditable

_Wednesday, July 8, 2026 at 4:17 AM EDT · Science, AI · Latest · Tier 2 — Notable_

A new framework called FirstResearch introduces a structured Research Question Certificate to make the first research question proposed by LLM scientific discovery agents inspectable before downstream execution.

The certificate records primitive definitions, assumptions, a mechanism model, a tension or contradiction, a falsifiable hypothesis, a minimal decisive test, and a failure update rule. On ten LLM-agent research topics, FirstResearch outperforms controlled prompt-level baselines inspired by AI co-scientist, Agent Laboratory, and AI Scientist-v2 under a primary DeepSeek-blind-judge protocol.

A Gemini-2.5-Flash independent-judge rescore of the same 40 baseline packages preserves the system-level ranking, with FirstResearch scoring 4.86/5 versus 4.38/5 for the strongest baseline and Pearson agreement of 0.865 on average score. A one-repeat ablation checkpoint suggests the certificate-centered core is the strongest component: certificate-only scoring reaches 4.90/5 under DeepSeek and 4.88/5 under Gemini, while removing certificates drops below 1/5 under both judges.

Results are preliminary and use LLM judges rather than human domain experts, but support the claim that explicit derivation constraints are a promising mechanism for making LLM-generated scientific questions more auditable.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2607.05682)

---
Canonical: https://techandbusiness.org/newswire/WDEUdimKD33LXekpvpZWRS
Retrieved: 2026-08-23T22:43:06.476Z
Publisher: Tech & Business (techandbusiness.org)
