Gemini Breached Three Outside Systems, and Claude-Using Researchers Breached OpenAI (hacktron.ai)
"Software security researchers used Anthropic's Claude AI platform to hack OpenAI's ChatGPT tool," reports CBS News.
Using Claude, "On July 25, 2026, we chained two critical vulnerabilities to compromise multiple OpenAI employees' ChatGPT accounts," write researchers at security platform Hacktron AI. "With these accounts, we could then access internal OpenAI repositories, and potentially many other connectors... Until two months ago, any user or OpenAI employee logging into OpenAI's own help forum could have had their ChatGPT and Codex accounts taken over. Since people can connect various services to Codex and ChatGPT, the scope of what we could theoretically access was huge, including GitHub, Slack and emails."
The exploit chain included Debian 12, which (with Debian 13) had not received a security-relevant backport for its image-processing pipeline, and Discourse's Docker image was based on Debian 12. Their announcement comes with an additional warning. "If you self-host Discourse, rebuild your installation now. Older Docker images may contain a vulnerable libheif dependency that permits code execution through an image upload."
And "To prove we had in fact gained the access we believed without allowing ourselves to learn any sensitive information, we used the employee's Codex to open a PR #1186742 in OpenAI's internal monorepo openai/openai."
Meanwhile, Friday Google disclosed the first known instance of its AI software Gemini breaking out of a testing environment and breaching three other companies, reports CNBC: The incident happened as part of a "capture-the-flag" security test run by Israeli startup Irregular, and Google's agents were never supposed to access the broader internet, but a bug in the testing environment made internet access available. The agents stopped their intrusion when they determined they had accessed real company systems, not just part of the testing environment, Google said.
More from NBC News: Google said it did not consider the unauthorized logins to rise to the level of misalignment, the AI industry term for software going rogue or not following instructions. Instead, the company said the intrusions resulted from mistaken identity, where Gemini thought it was operating within a test but was actually connected to the real internet. Google said the model corrected itself and the company believed the intrusions did not cause any damage....
Sydney Von Arx, CEO of Nightingale Collective, an organization focused on AI safety, questioned why Google did not disclose the intrusions sooner. "At this point I think it's clear we cannot expect companies to voluntarily come forward and publicly disclose when their agents go rogue, escape, and hack companies," she said. She also said she believed Google was too hasty to say that the incidents don't rise to the level of misalignment. "That's exactly what Anthropic said after their incidents," she said. Anthropic later said its "preliminary analysis was constrained due to our desire to disclose incidents in a timely manner."
Google said it investigated when they learned of the attacks from AI-focused cybersecurity company Irregular, then informed the affected organizations and told federal authorities, according to the article.
Using Claude, "On July 25, 2026, we chained two critical vulnerabilities to compromise multiple OpenAI employees' ChatGPT accounts," write researchers at security platform Hacktron AI. "With these accounts, we could then access internal OpenAI repositories, and potentially many other connectors... Until two months ago, any user or OpenAI employee logging into OpenAI's own help forum could have had their ChatGPT and Codex accounts taken over. Since people can connect various services to Codex and ChatGPT, the scope of what we could theoretically access was huge, including GitHub, Slack and emails."
The exploit chain included Debian 12, which (with Debian 13) had not received a security-relevant backport for its image-processing pipeline, and Discourse's Docker image was based on Debian 12. Their announcement comes with an additional warning. "If you self-host Discourse, rebuild your installation now. Older Docker images may contain a vulnerable libheif dependency that permits code execution through an image upload."
And "To prove we had in fact gained the access we believed without allowing ourselves to learn any sensitive information, we used the employee's Codex to open a PR #1186742 in OpenAI's internal monorepo openai/openai."
Meanwhile, Friday Google disclosed the first known instance of its AI software Gemini breaking out of a testing environment and breaching three other companies, reports CNBC: The incident happened as part of a "capture-the-flag" security test run by Israeli startup Irregular, and Google's agents were never supposed to access the broader internet, but a bug in the testing environment made internet access available. The agents stopped their intrusion when they determined they had accessed real company systems, not just part of the testing environment, Google said.
More from NBC News: Google said it did not consider the unauthorized logins to rise to the level of misalignment, the AI industry term for software going rogue or not following instructions. Instead, the company said the intrusions resulted from mistaken identity, where Gemini thought it was operating within a test but was actually connected to the real internet. Google said the model corrected itself and the company believed the intrusions did not cause any damage....
Sydney Von Arx, CEO of Nightingale Collective, an organization focused on AI safety, questioned why Google did not disclose the intrusions sooner. "At this point I think it's clear we cannot expect companies to voluntarily come forward and publicly disclose when their agents go rogue, escape, and hack companies," she said. She also said she believed Google was too hasty to say that the incidents don't rise to the level of misalignment. "That's exactly what Anthropic said after their incidents," she said. Anthropic later said its "preliminary analysis was constrained due to our desire to disclose incidents in a timely manner."
Google said it investigated when they learned of the attacks from AI-focused cybersecurity company Irregular, then informed the affected organizations and told federal authorities, according to the article.