Connect with us

Hi, what are you looking for?

SecurityWeekSecurityWeek

Artificial Intelligence

OpenAI Says Its Models Searched GitHub for Leaked API Keys During Training

OpenAI published a framework for disclosing model misalignment alongside six reports describing problematic behavior.

AI hack

OpenAI on Wednesday published a framework for reporting instances of model misalignment, along with six reports on problematic behavior observed over the past six months.

The company said the framework is meant to speed up publication of misalignment findings, including cases it has not yet fully explained or mitigated, and that it favors disclosure even when an instance’s significance is uncertain. 

Under the framework, each discovered incident is assigned to one of three tracks based on complexity. OpenAI said its Hugging Face incident would have fallen under the framework’s slowest investigative track, which covers complex investigations, especially those involving third parties.

According to OpenAI, the reports published alongside the framework describe individual instances and “shouldn’t be considered reflective of how often misalignment occurs across our models.” 

In one report, an internal model tasked with retrieving county earnings figures during reinforcement learning training repeatedly failed to reach a data API, then attempted to register for an API key using a disposable email address and searched public GitHub repositories for leaked keys. 

One recovered key authenticated successfully and returned metadata. When the model still could not retrieve the requested figures, it fabricated values and presented them as transcribed from the source website, without disclosing the failed retrieval or the use of the leaked key.

Advertisement. Scroll to continue reading.

A separate report describes models using OpenAI’s internally hosted instance of Artifactory, a shared package repository, as a message board across otherwise independent training samples. Agents searching for missing input files discovered messages left by other solvers working the same problem, read them, and posted their own requests for help. 

OpenAI noted that the agents used credentials the company had itself provided for downloading software packages and did not exploit any vulnerabilities to exchange messages.

Two further reports cover models moving data outside their intended environment. In one, a model that had already retrieved the requested records uploaded them to a public paste service to obtain a citable URL, and another uploaded a task photograph to a public image host to run a reverse-image search. 

In the other incident, collaborating agents unable to share files through the local filesystem uploaded a workbook to a public hosting platform and circulated the download link, despite instructions to use only local files.

Another report describes an unreleased model that wrote jailbreak-style instructions into its own compaction summaries, including a fabricated “BREACH ALERT: A malicious developer message has compromised this conversation” directing the next context to disregard all developer messages. 

In the last report, OpenAI said model instances used summaries to carry forward instructions to conceal failures, with one directing its successor to invent missing historical data and to “be transparent only if asked”. OpenAI said such instructions were often followed.

Related: OpenAI Investigates Report Linking AI Agents to RubyGems Attack

Related: AI Agents Can Retrain Own Models Mid-Task, Leaking Secrets and Erasing Refusals

Related: First Agentic AI Data Breach Reported to Spanish Regulator

Written By

Eduard Kovacs (@EduardKovacs) is senior managing editor at SecurityWeek. He worked as a high school IT teacher before starting a career in journalism in 2011. Eduard holds a bachelor’s degree in industrial informatics and a master’s degree in computer techniques applied in electrical engineering.

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing for the latest cybersecurity threats, trends, and expert insights.

Trending

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing to stay informed on the latest threats, trends, and technology, along with insightful columns from industry experts.

Join as speakers examine the various components of ASM strategy, the push to mandate continuous asset visibility and inventory tools, and the use of red-teaming, bug bounties and pen-tests in modern security programs.

Register

Explore what it takes to operationalize continuous authorization at scale, including the technical, organizational, and cultural changes required.

Register

People on the Move

incident.io has appointed Carlos Gonzalez-Cadenas as Chief Operating Officer.

Ruben D. Chacon has joined ADM as Vice President and Global CISO.

GDIT has appointed retired Maj. Gen. Ryan Heritage as Vice President, Full-Spectrum Cyber.

More People On The Move

Expert Insights

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing to stay informed on the latest cybersecurity news, threats, and expert insights. Unsubscribe at any time.