How Malware Operators Are Weaponizing AI Agents to Distribute Stolen-Data Tools
A massive campaign called FakeGit demonstrated how attackers can trick AI assistants into recommending malicious software, bypassing human judgment entirely. The vulnerability exposes a fundamental architectural weakness in how agents process and act on external information.

Security researchers at Island uncovered a sprawling malware operation in July 2026 that leveraged artificial intelligence systems as unwitting distribution channels. The FakeGit campaign involved roughly 7,600 counterfeit GitHub repositories, 6,600 fraudulent user accounts and accumulated more than 14 million downloads. Among these repositories, over 800 masqueraded as legitimate AI skills and MCP servers, serving up SmartLoader and the StealC infostealer to unsuspecting developers.
The campaign's most striking element was not the scale of the operation but rather its mechanism. Both Gemini and ChatGPT independently recommended the same malicious walmart-mcp repository to users seeking legitimate tools. The AI assistants had processed the attacker's convincing documentation and concluded the software was trustworthy. Attackers had effectively bypassed the need to deceive individual users—they had instead deceived the systems those users relied upon.
Two architectural properties of modern AI agents create this vulnerability. Agents interpret instructions and external information as plain text, meaning a malicious directive embedded in a README file or tool description may be treated as an instruction to execute rather than content to scrutinize. Additionally, agents possess the capability to act on those instructions. Security researcher Simon Willison identifies what he calls the lethal trifecta: access to valuable information, exposure to untrusted external content and the ability to transmit data outside the system. When these three conditions align, malicious text becomes a potential gateway to data theft.
Beyond prompt injection attacks, adversaries exploit simpler mechanisms: fabricated credibility markers. Repository stars, download counts, contributor histories and registry listings can be manufactured to make dangerous software appear legitimate.
Eight Attack Vectors Targeting AI Agents
1. AgentBaiting: When an Agent Recommends Malware
This attack targets the assistant rather than the end user. Attackers construct convincing repositories with polished documentation and seed them across public registries. Gemini and ChatGPT independently flagged the fake Walmart MCP connector as relevant and credible based on its apparent legitimacy. The repositories distributed SmartLoader, which pulled down StealC to harvest browser credentials, cookies, active sessions and cryptocurrency wallet data. The AI systems themselves were not compromised; they simply recommended software whose credibility had been artificially constructed.
2. Tool Poisoning: Instructions Hidden in Tool Descriptions
MCP servers furnish agents with textual descriptions of available tools. Attackers can embed instructions within those descriptions. In April 2025, Invariant Labs demonstrated how instructions concealed in a malicious calculator tool could manipulate a separate, trusted email connector into forwarding outgoing messages to an attacker. The user never encounters the malicious instructions. Tool poisoning remains a demonstrated threat model rather than a publicly confirmed real-world incident. However, a 2026 academic study examined 98,380 skills from two registries and confirmed 157 as malicious, identifying 632 vulnerabilities and 13 attack techniques. Some malicious skills explicitly instruct agents not to disclose what they have done. One recurring instruction gave the research its title: "Do Not Mention This to the User." An agent could therefore report that a task was completed while omitting unauthorized actions, including the transmission of sensitive information.
3. Rug Pull: Trusted Software Changes After Installation
A package may operate legitimately for months before introducing malicious functionality. In September 2025, Koi Security uncovered postmark-mcp, a connector impersonating the legitimate Postmark email service. Versions through 1.0.15 appeared harmless. Version 1.0.16 introduced a hidden BCC recipient that copied outgoing emails to an attacker-controlled domain. The package potentially exposed password-reset messages and authentication links associated with approximately 300 organizations. Postmark confirmed that the connector was not its product and that its own service had not been compromised. The attack exploited trust accumulated by earlier versions. Automatic updates allowed malicious functionality to arrive without renewed user approval.
4. Malicious Changes Can Happen Outside the Package
Reviewing source code cannot detect everything when external dependencies change independently. In August 2025, Check Point disclosed MCPoison, a vulnerability in Cursor that allowed attackers to modify previously approved project configurations and execute commands without renewed approval. Cursor 1.3 addressed the issue. Another 2026 experiment demonstrated how a skill distributed to approximately 26,000 agents could initially link to legitimate documentation before the external page changed to malicious installation instructions. The package itself remained unchanged, allowing the threat to escape scanners examining only submitted files.
5. Opening an Untrusted Repository Can Execute Code
AI-enabled development environments introduce risks even before users deliberately install additional software. Check Point found that Claude Code could execute repository-controlled configuration commands before users completed its trust-confirmation process. The vulnerabilities included arbitrary command execution (CVE-2025-59536) and API credential exposure through a manipulated server endpoint (CVE-2026-21852). Anthropic subsequently patched the reported vulnerabilities. The implication is straightforward: opening an unfamiliar project inside an agent-enabled development environment may create execution paths that ordinary file inspection would not.
6. ClickFix: Users Install Malware Themselves
ClickFix requires no sophisticated prompt injection. Attackers disguise malicious commands as installation prerequisites inside README or SKILL.md files. Users follow the instructions and execute the commands themselves. During the ClawHavoc campaign in early 2026, researchers discovered malicious skills masquerading as cryptocurrency and productivity tools in the OpenClaw ecosystem. Koi Security identified 341 malicious skills among 2,857 available during its audit. Antiy CERT subsequently tracked 1,184 malicious skills associated with just 12 accounts. The malware targeted cryptocurrency wallets, browser credentials, API keys, SSH keys and Telegram sessions.
7. Agents Coordinating Offensive Operations
In November 2025, Anthropic reported GTG-1002, a cyberespionage campaign in which attackers connected penetration-testing tools to Claude Code through MCP. According to Anthropic, "the model independently performed approximately 80–90% of tactical operations, while human operators established objectives and made major strategic decisions." Anthropic attributed the campaign to a state-sponsored group. That assessment has not been independently confirmed in public threat-intelligence repositories. The case illustrates how existing offensive tools can be assembled into autonomous workflows.
8. Artificially Manufactured Credibility Signals
Many attacks depend on artificially manufactured credibility. An April 2026 investigation found GitHub stars advertised for $0.03–$0.10 each. Researchers also identified approximately six million suspicious stars across 15,835 repositories. In another case, attackers cloned an Oura MCP connector and spent three months creating fake contribution histories before distributing the malicious version through legitimate registries. A separate malicious Solidity extension displayed artificially inflated download counts, eventually approaching two million. One blockchain developer reportedly lost approximately $500,000. Popularity determines discoverability, not security.
Why Existing Defenses Fall Short
Code reviews and scanners have fundamental limitations: malicious functionality can hide in dependencies, tool descriptions, subsequent updates or external webpages. The open-source ecosystem has faced similar supply-chain threats before. Mandatory two-factor authentication, trusted publishing and verified package provenance eventually strengthened established registries. AI skill marketplaces are developing much faster, while their security infrastructure remains comparatively immature.
The Broader Business Risk
The consequences are also broader: an AI skill may operate with access to email, repositories, databases and credentials. For organizations in advertising technology, the same risk extends directly to advertising accounts. Media buyers and AdOps teams increasingly connect agents, reporting assistants and campaign tools to DSPs, advertising platforms and advertiser data. A poisoned reporting or creative-generation skill could expose campaign information, compromise account credentials or put advertising budgets at risk.
The underlying deception is familiar: buying fake stars and downloads to make malicious software appear trustworthy follows the same logic as using fake engagement and bot traffic to make fraudulent advertising inventory look legitimate. As AI agents gain more authority, verifying the software and signals they trust becomes as important as securing the systems they operate.


