openai FrontierScience

Credit: OpenAI

OpenAI has unveiled FrontierScience, a new expert-level benchmark designed to measure how well AI models can perform advanced scientific reasoning—an ability the company believes is essential if AI is to meaningfully contribute to real research. Unlike traditional benchmarks that focus on factual recall or multiple-choice questions, FrontierScience targets the kind of deep, multi-step thinking scientists use to form hypotheses, test ideas, and synthesize knowledge across disciplines.

The announcement builds on a year of rapid progress in OpenAI’s most capable models. According to the company, systems such as GPT-5 have already demonstrated measurable impact in scientific workflows, helping researchers accelerate literature reviews, navigate cross-disciplinary research, and work through complex mathematical proofs. OpenAI’s November 2025 paper, Early science acceleration experiments with GPT-5, showed that these models can compress tasks that once took days or weeks into hours—suggesting AI is beginning to function as a practical research assistant rather than just an information retrieval tool.

FrontierScience was created to better measure this emerging capability. The benchmark includes more than 700 original questions written and verified by domain experts in physics, chemistry, and biology, with 160 forming a gold-standard evaluation set. It is divided into two tracks: FrontierScience-Olympiad, which contains 100 highly constrained problems designed by international science Olympiad medalists, and FrontierScience-Research, which features 60 open-ended, research-style subtasks graded using detailed rubrics. The Research track aims to reflect the type of reasoning a PhD-level scientist might encounter during real investigative work.

In early evaluations, OpenAI reports that GPT-5.2 leads other frontier models, scoring 77% on the Olympiad track and 25% on the Research track. While the Olympiad results suggest AI systems are approaching expert-level performance on structured scientific reasoning, the lower Research score highlights the difficulty of open-ended scientific thinking—where creativity, judgment, and experimental design remain challenging for current models. OpenAI says this mirrors how scientists already use AI today: as a tool to accelerate structured tasks, explore ideas, and surface connections, while relying on human expertise for framing problems and validating results.

The company is clear that FrontierScience is not a complete measure of scientific intelligence. It does not evaluate hypothesis generation, experimental execution, or multimodal research involving real-world lab systems. Still, OpenAI argues that the field urgently needs harder, more meaningful benchmarks to track progress as models improve. FrontierScience, they say, provides a sharper lens into where AI reasoning succeeds, where it fails, and what needs to improve if AI is to become a trusted collaborator in scientific discovery.

Looking ahead, OpenAI plans to expand FrontierScience to new domains and pair it with real-world evaluations that assess how AI systems contribute to actual scientific breakthroughs. Ultimately, the company says, the most important benchmark will be whether AI helps generate novel discoveries—but FrontierScience is a critical step toward understanding how close today’s models are to that goal.