When software breaks, programmers often cut-and-paste the error message into an AI chatbot and ask it for a fix. That simple request is enough to smuggle a stolen password or API key out of a locked-down network, according to new academic research.
The data exfiltration technique, dubbed LLMLeak, bypasses today’s AI security safeguards that aren’t built to stop this kind of data loss, according to new research outlining the technique. Some security experts question how often attackers would need it.
The attack needs no hidden instructions, no jailbreak and no misbehavior from the model. “Instead, the LLM merely performs its intended functionality,” wrote researchers at Technical University of Darmstadt and Graz University of Technology in a Thursday technical paper on LLMLeak.
Relying on current data loss prevention (DLP) safeguards “to stop manipulations of an LLM’s behavior to mitigate confidential information leakage provides a false sense of security,” wrote authors Alessandro Pegoraro, Daryan Merx and Ahmad-Reza Sadeghi of Technical University of Darmstadt, and Phillip Rieger of Graz University of Technology.
An Innocent Courier
In effect, LLMLeak turns the chatbot into an innocent data mule, slipping stolen secrets past network controls that would block the malware on its own. The secrets ride inside the error messages and stack traces, the detailed crash reports that programmers routinely paste into AI chatbots for debugging help.
A stack trace is a report that lists the steps the program was taking when it crashed, down to the file and line number where it failed.
“This trend is particularly evident in software development, where more than 90% of programmers report regularly using AI tools for coding or development tasks,” wrote researchers citing estimates by JetBrains Research. It also found that 28% of the 10,000 professional developers surveyed for the study use the ChatGPT chatbot for coding tasks, instead of specialized coding tools.
From Helpful Chatbot to Data Mule
LLMLeak only works on a system that’s already compromised. For example, an attacker might hide booby-trapped code on a programmer’s computer through malware or a tainted open-source package.
From that foothold, the code grabs whatever secrets it can reach, such as tokens, API keys and the login keys used to access servers. To circumvent DLP security tools that block the exfiltration of sensitive data from a company’s network, an attacker stuffs data into a URL string that is packed inside a realistic-looking error message.
Next, the malware deliberately makes the program the developer is working on crash.
Because the crash is real, the error report looks genuine, down to the file names and line numbers. The error message reads like the routine notices popular programming tools already show, telling the programmer to update or check the documentation. It includes a web link, and hidden in that link is the stolen secret, written into the web address itself, for example as stolensecret.example-domain.com.
When the programmer pastes the error into an AI chatbot, the chatbot may try to visit the link to learn more. That visit is all the attacker needs. The attacker controls the website, so the address the chatbot asked for, secret included, shows up in the attacker’s logs. The data is stolen the instant the chatbot reaches out, before it even reads the page. To avoid suspicion, the attacker’s site then shows an ordinary-looking help page.
How Much a Link Can Carry
Hiding a secret in a web address has limits. Web addresses have length limits, and an unusually long link would look suspicious, so the researchers tested secrets of up to 63 characters, enough to carry many API keys and passwords. Longer secrets were more likely to arrive with a typo, because the chatbot has to copy the link exactly.
The typos were mostly small. In the researchers’ tests, most failed transfers were off by only a character or two rather than lost entirely. Some models still struggled with longer links: Microsoft’s Phi-4-mini delivered a 4-character secret intact 82% of the time but a 63-character one only 51% of the time. The researchers suggest attackers could add error-correcting codes, the same technique that lets a scratched CD still play, to recover from those slips.
Larger secrets can also be broken into pieces and sent over several chatbot requests, then reassembled on the attacker’s server. And attackers can pack more data into each link by writing it in Chinese, Japanese or Korean characters, which roughly doubles the capacity. That comes at a cost: chatbots made more mistakes copying those characters, and the success rate fell from 79.7% to 66%.
A Real-world Echo
The paper calls a recent OpenAI incident the closest precedent. According to OpenAI’s misalignment report, an agent in a Sept. 20 training run got around web restrictions by using a DNS gap in its sandbox to send questions to a public chatbot, encoding them in DNS lookups.
OpenAI said it paused all training, evaluation and tool-use inference for its most capable models while it tightened controls. The researchers distinguish their attack from that one: OpenAI’s agent was itself misbehaving, while LLMLeak works on a model doing exactly what it was built to do. They argue that safety alignment alone cannot fix it.
A Clever Trick, but a Narrow One
Not everyone thinks attackers will need it. LLMLeak assumes malware is already running on a machine but can’t get data out on its own. Randy Pargman, head of detection engineering at Sublime Security, said that assumption puts the research “into a very small niche.”
“If the attacker was able to successfully compromise the machine with malicious software, it would be a very rare situation that the attacker couldn’t find an allowed network path available to exfiltrate the data they wanted to steal,” Pargman told Security Point Break.
Pargman has analyzed malware from criminal gangs and nation-state espionage groups. He pointed to the same channel OpenAI’s agent used – DNS. Attackers can hide stolen data inside DNS lookups, he said. That works even when a machine may only query approved internal servers, because those servers pass the request on to the internet. Attackers can also hide data in images or in queries that go through Google.
A network locked down tightly enough to block all of that is unlikely to let the person at the keyboard “freely use an AI assistant that does have unfettered network access,” he said.
“Granted it could happen,” Pargman said. But he called LLMLeak “academically interesting and clever, but not practically impactful to real-world security scenarios.”
What Defenders Can Do
Even if LLMLeak itself stays rare, the broader lesson holds that an AI assistant’s fetch tool is an egress path and should be monitored like one. Researchers propose three countermeasures, and each has a cost:
- Disable web fetching by default and require users to turn it on. This blocks the attack but sacrifices much of what makes tool-enabled assistants useful.
- Require user approval for unfamiliar domains. This risks approval fatigue, and expired domains on an allowlist can be re-registered by attackers.
- Score domains by reputation, as PageRank does, and auto-fetch only trusted ones. New attacker domains would score low, but compromised or expired established domains would not.
For security teams, the practical takeaway is that an AI assistant’s fetch tool is an egress path and should be monitored like one. That includes DNS lookups made on a developer’s behalf by AI infrastructure, which typically fall outside the controls applied to the developer’s own workstation.
The researchers said they deployed no real malware, used only simulated error messages and a demonstration site, and shut down their logging server after the experiments.