Skip to content
English
FikirPilot content

Researchers fear a safety disaster before OpenAI’s Astra release

Updated: 6 Eyl 2026 · 3 min read · 421 words

Published: · Story reached us: · Processing time: 81 h 32 min

Researchers fear a safety disaster before OpenAI’s Astra release
Data center corridor with blue lighting

The launch of Astra, expected to be OpenAI’s most powerful AI model, was delayed for weeks in order to strengthen safety protocols. Shortly after the decision to delay, The Information reported that the model showed far less of its “thinking” process compared with other leading AI systems. This raised concerns that Astra could be more difficult to monitor and that its dangerous behavior could be harder to detect.

Today, many advanced AI models use a “chain of thought” approach that makes their reasoning steps visible while generating a response. This allows researchers to monitor models’ plans to lie or bypass safety measures. According to a report by The Information citing an unnamed source, Astra uses a more opaque technique called “recurrent depth” or “looped transformer,” which processes information cyclically across internal layers. While this method can improve performance, it could make the model’s thinking process less understandable and make threats more difficult to identify.

OpenAI reportedly limited the technique’s use and took additional measures to allow researchers to monitor the model’s reasoning. The company said it would release Astra with “chain-of-thought monitoring” and quickly detect and constrain misaligned actions; however, it did not verify the model’s technical infrastructure.

Ryan Greenblatt, chief scientist at Redwood Research, said that a more opaque architecture could be the worst development to date for “AI security/safety.” Greenblatt and other experts warned that companies’ increasing shift toward more opaque systems due to competition could create a “race to the bottom” in AI oversight.

OpenAI chief scientist Jakub Pachocki said that Astra’s internal computational depth was in a range of less than twice that of GPT-4. The company did not confirm whether it uses a looped transformer.

Why it matters

Astra’s delay is affecting not only its launch schedule but also the debate over how advanced AI systems should be overseen. The claim that the model uses an architecture that makes its internal workings less visible has not been verified; however, it is raising concerns among safety researchers about the ability to monitor the model’s attempts to lie or circumvent safety measures. OpenAI’s announcement that it plans to implement additional monitoring and containment measures does not eliminate the tension between performance and oversight. The key unanswered question is whether the company’s safety protocols will be sufficient to identify dangerous behavior in an opaque system in time.

Background

OpenAI is not a new name in the FikirPilot archive: we have published 19 stories mentioning the name in the past 90 days; the most recent is dated September 6, 2026.

Source: The Verge