OpenAI Group PBC today started opening access to GPT-6 Astra, its newest and most capable large language model.
The company stated that the LLM demonstrates “state of the art” performance in multiple areas. The list includes coding, browsing and computer use, a term for tasks that require a model to interact with applications’ graphical interfaces. Furthermore, Astra has achieved perfect scores on three of the world’s toughest artificial intelligence benchmarks.
The first test that the model aced is called FrontierMath Tier 4. It comprises 50 math challenges that take most human mathematicians several weeks to solve. According to OpenAI, Astra completed the evaluation with a 98% score.
The model’s strong benchmark results aren’t too big of a surprise. Ahead of its public release, Astra solved several Erdos problems, decades-old math problems that are of particular interest to researchers. It also made several advances in a branch of computer science called computational complexity theory.
Astra scored 99.9% on a benchmark called ARC-AGI-3 that measures LLMs’ ability to learn new tasks. That capability is an important prerequisite to artificial general intelligence, or AGI. The model also achieved a perfect score on ExploitBench, which measures LLMs’ ability to find and exploit software vulnerabilities.
On Tuesday, OpenAI disclosed that Astra qualifies as a “critical” risk under its internal AI safety evaluation framework. An LLM receives that designation if it demonstrates the ability to hack “many well-protected systems” without human input. Astra’s cybersecurity capabilities required OpenAI to delay its release for several weeks. According to the company, its engineers spent that time developing guardrails against hacking.
OpenAI says that Astra’s ability to tackle complex tasks partly stems from a new data management approach.
LLMs generate a significant amount of data while generating prompt responses. When that data can’t fit in a model’s context window, a so-called compaction mechanism compresses it. Low-priority information is discarded to make room for new requests.
Astra also compresses prompts to make room for new data, but it doesn’t discard the original information. The model instead archives it in a searchable form that can be accessed later if necessary. That approach retains data points the model can use to improve its output quality.
At launch, Astra is only available to a limited number of customers via a program called Daybreak. The program enables organizations to use OpenAI’s models for cybersecurity research. The company will expand the availability of Astra over the coming days by bringing it to ChatGPT, Codex and its application programming interface.
OpenAI said today that it will provide Daybreak participants with $1 billion in credits. The goal is to help government agencies, utilities and other essential service operators improve their cybersecurity posture. The credits will cover not only model usage but also training and customer support. OpenAI will provide some of the training in partnership with MS-ISAC, a nonprofit that helps public sector organizations block cyberattacks.





