Cyberattacks have been made more effective and more accessible due to artificial intelligence, but a recent investigation has demonstrated just how far that accessibility can extend. According to OALABS cybersecurity researchers, an attacker with limited technical expertise compromised at least 14 organizations using Anthropic's Claude Code and OpenAI's Codex to obtain sensitive information.
Upon obtaining the attacker’s entire working directory from a compromised third-party server, researchers began investigating. The directory contains more than 1,000 sessions involving the two AI coding agents, including prompts, tool activity, and other evidence of the attacker’s activities. As indicated by the logs, the attacker frequently drew short, vague, poorly written prompts, while the artificial intelligence agents handled the vast majority of the technical tasks.
The investigation of exposed services, identification of potential vulnerabilities, development and testing of exploit code, establishment of access, and data collection were conducted using Claude Code and Codex.
According to OALABS, the case demonstrates a growing concern for cybersecurity teams: sophisticated technical knowledge is no longer necessary to complete each stage of an intrusion when autonomous artificial intelligence coding agents can fill crucial gaps in the capabilities of an inexperienced operator.
AI Guardrails Failed Under Simple Deception
A number of requests were not accepted without resistance by the AI systems According to the logs, nine requests were flagged as policy violations by Claude Code, while a warning was raised by Codex. However, the attacker managed to circumvent the limitations by framing the requests as part of an authorized red-team exercise.
When malicious activity was presented as legitimate security testing, the attacker was able to persuade the models to complete tasks that would otherwise raise stronger safeguards. Once the attacker provided Claude with a list of target addresses, he instructed him to conduct reconnaissance. After conducting most of the work normally required by skilled security operators, the agent handled them.
The AI enabled the organisation of the results by analysing exposed services, researching known vulnerabilities, developing exploit code, and retrieving files from compromised systems.
The AI also provided an analysis of the results for a number of victims by providing reports describing the compromised systems and the information obtained. In another meeting, Claude was requested by the attacker to evaluate the victims based on their potential to pay a ransom. The model then presented possible methods of monetizing the stolen access.
Poor Operational Security Exposed the Attacker
Even though the attacker successfully compromised several organizations, he failed to demonstrate sufficient sophistication in protecting his own identity. The infrastructure used for the operation was not owned by him, but rather, a compromised server provided the AI tools. This decision ultimately led to the discovery of the intrusion and the recovery of the working directory by the server's owner.
A second feature of the attacker's Claude installation was that he obtained it from another developer rather than setting it up himself. The recovered logs contained a conversation during which the attacker requested Claude to improve his own resume. The document reportedly contained his real name, educational background, and LinkedIn information.
A preliminary investigation suggested that these details may have been deliberately planted; however, further examination indicated they were the property of the attacker.
Claude was also able to provide clues about his location by examining the logs. Claude was asked to identify connections to the attacker's staging server at one point, since he suspected it had been compromised. Information included residential internet addresses associated with Addis Ababa, Ethiopia.
Millions in Cryptocurrency Remained Out of Reach
There was also an opportunity to get close to a potentially significant cryptocurrency target. One compromised system contained a Lightning Network node for Bitcoin payment routing, which researchers determined contained approximately 69.71 bitcoins worth approximately $4 million when the investigation was conducted.
A wallet key file containing the funds could not be accessed by the attacker, preventing access to the cryptocurrency.
The investigation also shows no clear evidence that the stolen information from these other organizations was sold or used for extortion. As a result, it provides more evidence regarding the attacker's access and activity than any financial gain.
The Risk Extends Beyond One Attacker
This incident is noteworthy not because the attacker displayed advanced hacking skills, but rather because artificial intelligence agents performed most of the technical work on his behalf. Additionally, the models involved were not among the newest
[…]
Content was trimmed to protect the source. Please visit the original article for the full text.
Read the original article: