
AI Benchmark Tests Real Scientific Discovery Skills
A new benchmark called TRACES evaluates whether AI can make genuine scientific discoveries instead of just recycling existing knowledge. Unlike traditional tests with preset answers, it places AI systems in real-world problem-solving environments where they must adapt, experiment, and navigate uncertainty.
Artificial intelligence may soon prove whether it can truly discover new scientific knowledge, not just repeat what it already knows.
Apodex has launched TRACES, a groundbreaking benchmark that tests AI systems on open-ended scientific problems where the correct answer isn't yet known. This marks a major shift from conventional AI testing, which typically relies on static datasets and predetermined answer keys.
The benchmark places AI systems inside executable environments that mirror real scientific work. Depending on the problem, these environments can include scientific literature, datasets, code execution, specialized tools, simulators, and experimental feedback. The AI must observe, act, use tools, receive feedback, and revise its approach while working toward a verifiable outcome.
Dr. Sheng Wang, Lead Scientist at Apodex, explains that TRACES evaluates what the company calls "discoverative AI." These are systems designed to identify genuinely new findings from existing knowledge rather than reproduce information already contained in their training data.
The benchmark scores AI across six key capabilities. Tools measures whether the system selects and correctly uses external resources. Repair evaluates its ability to identify and fix errors after receiving feedback. Alternatives checks whether competing hypotheses are considered and revised as evidence accumulates.

The remaining three capabilities focus on scientific rigor. Coherence evaluates whether the system maintains logical consistency across longer problem-solving sequences. Evidence assesses whether conclusions are grounded in actual observations, data, experiments, or citations. Scope examines whether the system accurately defines where a conclusion applies and where it doesn't.
Why This Inspires
What makes TRACES revolutionary is its focus on the journey, not just the destination. Brian Wang, AI Research Scientist at Apodex, notes that in scientific discovery, the final answer represents just one line at the end of hundreds of judgments about what to try next, when evidence is sufficient, and when to abandon a hypothesis.
The benchmark combines outcome and process verification. An outcome verifier grades submissions against hidden ground truth where available, while a process verifier evaluates whether the reasoning and evidence supporting the conclusion meet scientific standards. An independent model reviews each evaluation, with disagreements triggering re-scoring and adjudication.
Apodex has assembled 423 high-value problems from a survey spanning 561 industries across 16 sectors. The benchmark is open to teams developing models, agent systems, and solver frameworks. Researchers and organizations can also submit scientific problems to be converted into executable evaluation environments.
The stakes are high because distinguishing between AI that reproduces knowledge and AI that can work through uncertainty could accelerate discoveries in medicine, climate science, and countless other fields. If AI can truly navigate the messy, uncertain process of real scientific discovery, it could become a genuine partner in solving humanity's most pressing challenges.
Based on reporting by Google: scientific discovery
This story was written by BrightWire based on verified news reports.
Spread the positivity!
Share this good news with someone who needs it

