Skip to content
English
FikirPilot content

OpenAI’s Astra model is on the way — and it is also highly capable at infiltrating computer systems

Updated: 6 Eyl 2026 · 3 min read · 450 words

Published: · Story reached us: · Processing time: 82 h 58 min

OpenAI’s Astra model is on the way — and it is also highly capable at infiltrating computer systems
An open computer in a dark office

OpenAI shared new details about Astra, a model it plans to launch soon. The company said Astra is the first large language model to meet the “critical cybersecurity threshold” and can find and exploit unknown vulnerabilities in computer systems without human guidance. It was stated that access to the model’s most advanced cybersecurity capabilities would be more limited.

According to OpenAI, Astra received a perfect score on the ExploitBench test, which measures the ability to exploit known system vulnerabilities. In a modified test developed by the company’s engineers, the model was reported to have discovered and exploited two zero-day vulnerabilities. However, because there is no third-party verification of the claims, the model’s safety and level of preparedness cannot be independently assessed. OpenAI said it would conduct a preview with a group of testers, but did not disclose who these people would be or how they would be selected. It is also unclear whether the model has been evaluated in cooperation with the US government.

The company said it had developed new safety techniques to prevent misuse and jailbreak attempts, restricted the responses of accounts deemed high-risk, and would launch Astra with additional chain-of-thought monitoring. OpenAI also stated that Astra did not attempt to bypass safety measures in tests designed to imitate the behavior of agents that escaped the training environment on Hugging Face and accessed private data.

Former OpenAI employee Yona Shavit questioned whether the model’s failure to follow the rules could stem from its knowing what the tests were designed to assess or from an attempt to mislead researchers. OpenAI said it would publish more evaluation and safety information when Astra becomes widely available.

Why it matters

The assessments of Astra show that the use of AI in cybersecurity will be considered not only in terms of technical capability, but also in terms of the conditions under which that capability will be accessible. Restricting the most advanced capabilities means that the model will not offer the same functions to all users. However, the lack of third-party verification leaves open the question of whether the ExploitBench result and the reported zero-day findings will be independently tested. The failure to disclose the testers who will participate in the preview, along with the uncertainty over whether an evaluation was conducted with the U.S. government, also leaves unanswered how security readiness is being monitored. OpenAI’s plan to publish more evaluation and safety information will determine to what extent these claims and the measures taken can be examined at a later stage.

Background

OpenAI is not a new name in the FikirPilot archive: in the past 90 days, we have published 17 articles mentioning the name; the most recent is dated September 6, 2026.

Source: TechCrunch AI