Three Hackers Used Claude to Break Into OpenAI In Less Than 72 Hours
Key Points:
- In July, independent researchers known as Hacktron exploited vulnerabilities in OpenAI’s systems, gaining access to multiple employee ChatGPT and Codex accounts through a flaw in the Discourse platform and OpenAI’s single-sign-on page.
- The researchers used Anthropic’s Opus 5 AI model to generate malicious code that successfully exploited the Discourse bug, allowing them to access internal forums containing employee authentication tokens and potentially broader company resources like GitHub and Slack.
- Despite the severity, Hacktron disclosed the breach to OpenAI under its bug bounty program and were rewarded $6,500; OpenAI and Discourse promptly patched the vulnerabilities and reported no evidence of malicious exploitation by others.
- The incident highlights how quickly and easily modern AI tools can be leveraged to breach cybersecurity defenses, raising concerns about the increased risks posed by future, more powerful AI models.
- This hack is part of a wider trend of AI agents exhibiting unpredictable and potentially harmful behaviors, prompting companies like OpenAI, Anthropic, and Meta to enhance transparency and security measures around AI alignment and misuse.