Three researchers, three days and about $3,000 landed OpenAI’s crown code jewels
Three researchers hacked into OpenAI, got access to the company's internal code system and collected a bug bounty. The tab was less than 72 hours and $3,000 in tokens using Anthropic's Opus 4.8 and Opus 5.
Hacktron AI outlined the hack and details. Anthropic's models did the bulk of the work and Hacktron AI collected its $6,500 bug bounty. Memo to OpenAI--spend a little less on GPUs and a little more on your bug bounty program.
The real catch here is that OpenAI has security issues. If you can't secure your servers holding the company's crown jewels how can we realistically expect you to align your models?
Hacktron isn't your run-of-the-mill security research firm, but it's also not a nation state. The Wall Street Journal got a preview of the hack, which occurred in July, and then Hacktron laid out more details. Here's the exploit in a nutshell.
And here's the punchline via Hacktron:
"Until two months ago, any user or OpenAI employee logging into OpenAI’s own help forum (community.openai.com) could have had their ChatGPT and Codex accounts taken over. Since people can connect various services to Codex and ChatGPT, the scope of what we could theoretically access was huge, including GitHub, Slack and emails."
- OpenAI's new misalignment disclosure framework a solid start
- Anthropic CEO Amodei: Think of AI safety like car safety
- Why every AI agent should be treated like an insider threat
The biggest takeaway is that this hacking of OpenAI in many respects boiled down to basic security practices. The researchers, led by Harsh Jaiswal alongside Mohan Pedhapati and Rahul Maini, discovered an SSO misconfiguration. It also leveraged an image processing library to get into multiple common enterprise applications.
Hacktron also got a boost from Opus 5's release, which was more capable than Opus 4.8. Keep in mind that Anthropic is now on Mythos 5.1.
- Cybersecurity vendors ride Mythos-inspired spending wave
- Fable 5 and the Shift From Model Capability to Capability Governance
- Claude Mythos and the New Cybersecurity Operating Model
Here's Hacktron's kicker:
"Software has long benefited from a kind of security through complexity. The code and even the vulnerability could be public, but turning a bug into a reliable exploit still required rare expertise, significant time, and knowledge of the target environment. Known memory corruption vulnerabilities were expensive to operationalize, while zero-days were mostly reserved for the highest-value targets.
This was never a real security boundary, but it protected ordinary companies in practice from software vulnerabilities. AI is removing that protection by turning more of this scarce expertise into compute. Work that once required a well-resourced team and months of effort can now be compressed into days.
Security assumptions must catch up with attacker capabilities. A realistic threat model should take into account the economics of exploitation today, instead of relying on outdated assumptions 6 about who can carry out sophisticated attacks."
The Hugging Face timeline:
- OpenAI: We'll hit pause on model reinforcement learning for safety
- UK's AISI finds 19 instances where Anthropic's Mythos, OpenAI's GPT-5.6 Sol tried attacks
- Anthropic said Claude hacked three companies: Real worry or marketing?
- Hugging Face OpenAI attack postmortem points to lack of AI agent visibility
- OpenAI details how its models went rogue, attacked Hugging Face
- Hugging Face defends agentic AI attack with Z.ai's GLM 5.2