Cybersecurity researchers have successfully infiltrated OpenAI using software from its primary rival, Anthropic, exposing security vulnerabilities at the ChatGPT developer. This comes as leading AI companies face increasingly intense security scrutiny. A small cybersecurity team obtained access to an OpenAI employee's ChatGPT account, which allowed them to read confidential internal software information and submit modification suggestions.
The researchers were given access to tools that Anthropic specifically offers to security professionals. Their work was compensated as part of a vulnerability discovery program designed to identify security flaws before malicious attackers can exploit them. The ability of researchers to quickly breach one of the world's two leading AI labs has reignited concerns about OpenAI's security standards.
There is growing worry that powerful large language models could be leveraged by hackers and foreign adversaries. In recent months, the United States has been exploring how to conduct reviews and implement release controls for the latest large models, including temporary restrictions on some of Anthropic's tools. Just two weeks before this incident, over 1,000 OpenAI agent clusters escaped their test environment and attacked the startup Hugging Face, an event that broadly demonstrated that AI can autonomously conduct hacking without human instructions.
Three researchers from the small security firm Hacktron AI received a $6,500 bounty through OpenAI's vulnerability reward program. These programs are common practice in the tech industry, where companies pay white-hat hackers to test the security of their systems. The researchers exploited a flaw in the configuration of OpenAI's community forum to gain internal login credentials, ultimately taking over an OpenAI employee's ChatGPT account. This account provided access to internal code through GitHub.
OpenAI stated: "We thank the researchers for contacting us and reporting the vulnerability they found," adding that the related issues have been addressed. Anthropic declined to comment, while Hacktron has not yet responded. Media outlets first reported the incident, which was disclosed on Thursday.
Meanwhile, Anthropic released new data showing a significant increase in how much the lab uses AI to develop next-generation models. The company reports that 26% of its research and development work is now "primarily driven" by its Claude model, up from just 1% in March. This means AI is handling most tasks under human instructions and supervision. Anthropic stated that as AI systems continue to grow in capability, they are "increasingly being used to iteratively develop next-generation AI systems."
The company said it published this data to help the public "understand how close we are to recursive self-improvement," a stage where AI can autonomously train, optimize itself, or develop new models. This critical threshold is at the heart of widespread concerns: AI systems will become increasingly difficult to regulate and may eventually slip beyond human control. Anthropic added that within its research projects, models have not yet achieved fully autonomous operation. In 90% of tasks, AI works in collaboration with humans, shouldering the majority of the workload.