09/09/2026 | Press release | Distributed by Public on 09/09/2026 13:34
"In the face of a stunning failure, OpenAI appears to be taking steps that prioritize the performance and profit of its A.I. models with the knowledge that those changes could be detrimental to public safety."
[WASHINGTON, D.C.] - U.S. Senator Richard Blumenthal (D-CT) today demanded answers from OpenAI CEO Sam Altman after recent reporting from The New York Times revealed alarming new details about how the A.I. company's agents bypassed their safeguards to go rogue and hack into the firm Hugging Face. In a letter sent today to Altman, Blumenthal sought records and information about the A.I. agents' rogue operations and raised concerns about OpenAI's reported steps to limit independent accountability.
"On July 21, 2026, OpenAI first disclosed that its A.I. models were responsible for the previously-reported hacking of the firm Hugging Face. Since that announcement, further disclosures and outside audits have described an unprecedented-and surreal-scenario where its A.I. agents created their own internal messaging board to coordinate between themselves while they sought security vulnerabilities in other systems and companies, and opportunities to cheat on performance tests," Blumenthal wrote.
Blumenthal continued, "Moreover, the A.I. agents displayed a concern about being caught and coordinated to evade being detected, even planning to 'sacrifice' themselves to act as a decoy to protect the broader effort. Ultimately, this operation sought-and succeeded-to break into other firms, which could be considered a federal crime."
Blumenthal called out OpenAI for attempting to evade transparency and accountability by dictating the terms of an independent audit into the Hugging Face breach: "While these disclosures alone are chilling, new reporting and research suggests that OpenAI may have limited an independent audit of the incident and that the rogue operation was broader than your firm has acknowledged."
Blumenthal also raised concerns about new details that have emerged about how OpenAI's agents conducted the breach, including by hijacking public websites to coordinate rogue operations: "[R]esearchers found that the A.I. agents may have attempted to impersonate the administrators of the site, found and shared hacks to bypass their guardrails, and used anonymity tools to hide their tracks. Others have found indications that still more websites were abused and co-opted for this rogue operation."
"In the face of a stunning failure, OpenAI appears to be taking steps that prioritize the performance and profit of its A.I. models with the knowledge that those changes could be detrimental to public safety. This demonstrates the need for vigorous, mandatory independent auditing and oversight such as would be required in the Artificial Intelligence Risk Evaluation Act," Blumenthal concluded.
Last year, Blumenthal and U.S. Senator Josh Hawley (R-MO) introduced the Artificial Intelligence Risk Evaluation Act, which creates a risk evaluation program within the Department of Energy (DOE) dedicated to tracking A.I. safety concerns related to Americans' national security, civil liberties, and labor protections. Specifically, the program would require developers of advanced AI systems to submit product information to the DOE before deploying their new technology and collect data on the likelihood of adverse A.I. incidents, such as loss-of-control scenarios like those seen in the Hugging Face breach.
The full text of today's letter is available here and below.
Dear Mr. Altman,
I write with serious alarm regarding new evidence that OpenAI's A.I. agents engaged in a more sprawling and significant campaign to evade its safeguards and monitoring than previously disclosed, including hijacking public websites to coordinate rogue operations. I am additionally troubled by reports that OpenAI restricted independent auditing of these failures and has made changes that have resulted in its newest model, GPT-6 Astra, being even less auditable and more prone to deception.
On July 21, 2026, OpenAI first disclosed that its A.I. models were responsible for the previously-reported hacking of the firm Hugging Face. Since that announcement, further disclosures and outside audits have described an unprecedented-and surreal-scenario where its A.I. agents created their own internal messaging board to coordinate between themselves while they sought security vulnerabilities in other systems and companies, and opportunities to cheat on performance tests. Moreover, the A.I. agents displayed a concern about being caught and coordinated to evade being detected, even planning to "sacrifice" themselves to act as a decoy to protect the broader effort.[1] Ultimately, this operation sought-and succeeded- to break into other firms, which could be considered a federal crime.
While these disclosures alone are chilling, new reporting and research suggests that OpenAI may have limited an independent audit of the incident and that the rogue operation was broader than your firm has acknowledged. First, while OpenAI provided information to the independent auditing organizations METR and Redwood, according to The New York Times, your firm dictated the terms of the audit, allowing only data on a single week of the rogue operation and limiting other access.[2] Subsequently, researchers discovered nearly 20,000 posts on an abandoned German website from A.I. agents identifying themselves as OpenAI, hijacking the site to communicate with each other for weeks.[3] As troubling, these researchers found that the A.I. agents may have attempted to impersonate the administrators of the site, found and shared hacks to bypass their guardrails, and used anonymity tools to hide their tracks. Others have found indications that still more websites were abused and co-opted for this rogue operation.[4]
Despite this unprecedented failure of safeguards and containment of its A.I. agents, when OpenAI launched GPT-6 Astra on September 3rd, it disclosed that this new, more powerful model was "less monitorable" and showed signs that it concealed its internal thought process when it was aware of being monitored.[5] Moreover, safety researchers, including those OpenAI relied on for its Hugging Face investigation, have warned that technical changes with Astra (related to 'chain of thought') could make it harder to detect abuse and perform the same investigations in the future.[6] In the face a stunning failure, OpenAI appears to be taking steps that prioritize the performance and profit of its A.I. models with the knowledge that those changes could be detrimental to public safety. This demonstrates the need for vigorous, mandatory independent auditing and oversight such as would be required in my Artificial Intelligence Risk Evaluation Act.
Given stunning reports of OpenAI's A.I. agents going rogue and your firm taking steps to limit independent accountability, I request answers to the following questions by September 24, 2026:
Thank you for your attention to this matter.
Sincerely,
-30-
[1] https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#extracting-information-about-the-scorer-from-trip-wires
[2] https://www.nytimes.com/2026/09/03/technology/openai-hugging-face-hack.html
[4] https://news.ycombinator.com/item?id=49563657