OpenAI’s new reasoning technique alarms AI safety experts
Key Points:
- OpenAI’s new Astra model employs a reasoning technique called “recurrent depth” or “opaque recurrence,” which processes queries in loops rather than sequential steps, potentially making its chain of thought harder to monitor.
- AI safety experts, including Redwood CEO Buck Shlegeris and advocate Zvi Mowshowitz, have expressed serious concerns that this technique could undermine the ability to track and ensure safe AI behavior, warning it may lead to a “race to the bottom” in AI transparency.
- Despite these worries, Astra’s use of opaque recurrence is reportedly limited, with OpenAI maintaining its commitment to legible chain-of-thought monitoring and rejecting any move toward unintelligible “neuralese” reasoning.
- The technique’s emergence has prompted discussions at other AI labs like Anthropic and Google DeepMind, raising broader industry concerns about the potential for reasoning to shift into latent spaces that are less interpretable.
- Redwood Research’s Ryan Greenblatt highlighted the risk that scaling opaque reasoning could eventually remove visible reasoning traces altogether, urging caution and hoping OpenAI will halt further development of such architectures.