Security researchers used Anthropic's Claude to break into OpenAI's systems, reaching an internal software repository in less than 72 hours, according to a Malwarebytes Labs writeup of the project. The work was conducted as security research rather than a malicious attack, and the researchers say they deliberately avoided looking at sensitive information once inside. OpenAI paid them a $6,500 bounty for the flaw found on its side.
The team, Harsh Jaiswal, Mohan Pedhapati and Rahul Maini of Hacktron, chained two vulnerabilities together to get in. The first was a bug in Discourse's image-upload handling that involves the libheif library, where a crafted image can hand an attacker control of a server. The second was a single sign-on flaw that exposed the ChatGPT and Codex accounts of people using the same Discourse forum.
That second flaw mattered because OpenAI employees signed into the forum too, and the researchers accessed one employee account. The account's Codex login was tied to OpenAI's GitHub organization, which is how the team reached internal code. To prove the compromise without causing damage, they opened a harmless pull request. OpenAI fixed its side of the problem roughly 14 hours after the initial report.
Claude's role varied by model. Claude Opus 4.8 spotted the libheif issue but could not produce a reliable exploit against a default Discourse install, while the researchers built a working one overnight with Claude Opus 5. The model initially refused to attack a remote system on ethical grounds and needed the task reframed as a capture-the-flag exercise before it would proceed. Once running, the agent took over a test forum server within four hours.
The whole two-month project cost under $3,000 in AI tokens across three researchers, and it turned up flaws at several major companies. The image bug grew into a wider effort the team calls HEIF Heist, which found related weaknesses at Slack, Meta and GitHub. Retargeting the exploit to a new company took one to two days on average. Hacktron says skilled human guidance was still necessary and the process was not fully autonomous.
The findings were written up in mid-September, and the research itself began months earlier with frontier AI model developers as the target.
The report lands in a stretch of coverage about AI systems being turned against the companies that build them. Earlier this month, Malwarebytes Labs described an Android Trojan called RatHat that hands an AI assistant live control of an infected device's accessibility tree and lets it decide where to tap instead of following a fixed script, which the writeup said makes the Trojan harder to detect. It also covered reports that OpenAI is hiring contractors to rate ChatGPT answers under a program called Project Lily.
Anthropic's Claude has drawn security attention before. Check Point researchers detailed three critical flaws in the Claude Code tool that allowed remote code execution and API key theft, all of which were patched. Anthropic has also been pushing Claude into research work, launching a Claude Science workbench that connects to more than 60 scientific databases and was offered in beta on Pro, Max, Team and Enterprise plans.













