
GeneBench-Pro is a research-level benchmark that evaluates AI agents' ability to make complex, judgment-based decisions in computational biology, focusing on real-world challenges like ambiguity and data analysis. It includes 129 problems across 10 domains and 21 sub-domains, with a focus on system-level reasoning and decision-making. The benchmark assesses progress in overcoming computational biology's challenges, with frontier models like GPT-5.6 Sol achieving a pass rate of 28.7%, highlighting the difficulty of the tasks and the potential for partial automation.
See how RingCentral uses ChatGPT Work and Codex to accelerate AI product development and centralize operational intelligence across engineering and operations.

OpenAI and AWS are making Daybreak cybersecurity capabilities available through Amazon Bedrock to support enterprise security workflows.

OpenAI begins testing ads in ChatGPT to support free access, with clear labeling, answer independence, strong privacy protections, and user control.
