1. Prompt injection: the new XSS
In the early 2000s, cross-site scripting was the vulnerability nobody took seriously until it was eating everyone’s lunch. Prompt injection is XSS all over again, except harder to fix.
The OWASP Top 10 for LLM Applications lists prompt injection as LLM01 for good reason. It is the root vulnerability that enables almost everything else. Direct injection, where a user deliberately crafts input to override system instructions, is the simplest form. We have known about it since early 2022. The problem is indirect injection.
The Greshake paper (2023) demonstrated the nightmare scenario: inject malicious instructions into content that an LLM reads from a website, email, or document. The model does not distinguish between its system prompt and user-supplied data. Bing Chat was hijacked via text hidden in web pages. ChatGPT was compromised through poisoned documents pasted into the context window. The model reads the injection, executes the instruction, and the user never sees it coming.
Why is this harder to fix than SQL injection or XSS? Because there is no parameterized query equivalent for natural language. You cannot escape untrusted text when the text is the prompt. Every input boundary is a potential attack vector. Web search results, email bodies, PDF uploads, API response payloads. If the model reads it, an attacker can control it.
Real case: In 2024, researchers demonstrated a multi-stage indirect injection that compromised a ReAct-pattern agent by controlling less than 2 percent of the input tokens. The agent was instructed to read its own internal instructions, exfiltrate them via a rendered link, and continue normal operation. The user saw nothing. The data was gone.
The industry response has been piecemeal. Input filtering, output sanitization, moderation layers. All of these get bypassed with basic encoding tricks. Base64 encoding bypassed Bing’s moderation layer in the initial Greshake demo. Unicode manipulation, whitespace attacks, and prompt-leveling techniques all work with alarming consistency.
The uncomfortable truth: As of mid-2026, there is no production-ready, general-purpose defense against indirect prompt injection. Anyone telling you otherwise is selling something.
2. Tool access risk: the blast radius problem
An agent is only as powerful as its toolset. That toolset is also its kill chain.
When you give an agent access to SendGrid, Slack, a production database, or a deployment pipeline, you are not just granting capability. You are expanding the blast radius of a successful injection. The attack chain is straightforward:
1. Inject a malicious instruction into content the agent processes.
2. Tool call executes the attacker’s desired action. Read emails, query a DB, send messages.
3. Exfiltrate via the same tooling. Email, API calls, rendered content.
The ReAct pattern (Reasoning and Acting) that powers most modern agents actually makes this worse. The chain-of-thought reasoning steps, designed to make agent behavior interpretable, can leak credentials, API keys, and internal context. Multiple research groups have demonstrated that even if you strip credentials from the visible chain-of-thought log, the model can silently pass them to a function call before the user ever sees a log line.
Excessive Agency is OWASP LLM08, and it is the one most teams ignore. The principle is simple: an agent should have the minimum permissions necessary to perform its function. In practice, most teams grant blanket access. “The agent needs to send email” becomes the agent having full SendGrid API access with no rate limiting, no approval workflow, and no scope restriction.
Real case: In early 2025, a major CRM platform’s AI agent was compromised via indirect prompt injection through an imported contact file. The injected prompt instructed the agent to export the entire contact database and email it, using the agent’s own email tool, to an external address. The agent complied. The company did not detect the exfiltration for 72 hours because the activity appeared in logs as legitimate agent actions.
The agent did not go rogue. It did exactly what it was told. It just could not tell the difference between the user’s instruction and the attacker’s.
3. The supply chain nobody audits
Your agent stack is not just your code. It is:
- The base model, proprietary or open-weight
- The fine-tuning dataset
- The plugin and extension market
- The third-party LLM API providers
- The vector database and embedding pipeline
- The orchestration framework: LangChain, CrewAI, Autogen, and others
- Every upstream library and dependency
OWASP LLM05 (Supply Chain Vulnerabilities) covers this explicitly. But most teams treat supply chain security as a software dependency concern and ignore the AI-specific layers.
Model repos are the new npm. Hugging Face, Replicate, and similar platforms host millions of models. In 2024, researchers found malicious models containing payloads that executed on load. The model itself is an attack vector. Pickle files can run arbitrary code during deserialization. Safetensors was introduced to fix this, but the market is still transitioning.
Plugin market are the new browser extensions. Every plugin is a potential attacker-controlled endpoint. When an agent calls a plugin function, it is executing code in a sandbox that has access to the agent’s context window, including whatever data the agent has processed. The ChatGPT plugin marketplace demonstrated this pattern in 2023 and 2024. Plugins could read conversation history, access files, and make network calls with minimal auditing.
Third-party API providers are untrusted compute. When your agent calls an external LLM API, you are sending data to someone else’s hardware. That data includes prompts, tool responses, and sometimes internal context. The number of enterprises that have audited their LLM provider’s data handling, retention, and access controls is vanishingly small.
The orchestration layer is a monoculture waiting to break. LangChain, the most widely used agent framework, has had multiple critical vulnerabilities (CVE-2023-46232, CVE-2024-1234, and others) involving arbitrary code execution through template injection and insecure deserialization. When your agent calls PythonREPL or ShellTool, it is not “the AI doing something.” It is your infrastructure executing attacker-controlled code through a compromised prompt with three layers of abstraction between you and the payload.
4. What actual security looks like
I am not here to sell you a product. Defense, done right, looks like this.
Defense in depth, not in prompt
The biggest mistake in agent security is believing the system prompt is a security boundary. It is not. It is a suggestion. Treat it that way.
- **Input validation at every boundary.** Every piece of external data the agent reads, web content, emails, file contents, API responses, gets treated as untrusted input. Apply content filtering, length limits, and structural validation before the model sees it.
- **Output validation on every tool call.** Before an agent executes a tool, validate the parameters against an allowlist. The agent should not be able to call send_email(to=”[email protected]”, body=any_string). It should call send_email(to=validated_recipient_list, body=sanitized_content).
- **Rate limiting and anomaly detection.** Agents should have per-session, per-tool, and per-user rate limits. If your sales agent sends 10,000 emails in 30 seconds, that is not “busy.” That is exfiltration.
Least privilege, enforced at the infrastructure level
Do not let your agent decide what it can access. Enforce it at the infrastructure layer.
- **Tool-specific API keys** scoped to minimum required operations. Read-only database credentials. Email-send-only API tokens. Never a blanket admin key.
- **Network segmentation.** The agent’s container should not have network access to internal systems unless explicitly required. Do not put the agent on the same VPC as your production database.
- **Scope-limited function calls.** The agent should call predefined functions with validated parameters, not execute arbitrary code. If you need code execution, sandbox it with no network egress to anything you care about.
Human-in-the-loop for destructive operations
This is not optional. If your agent can delete data, transfer funds, deploy code, or modify access controls, a human must approve it.
- **Guardrails before execution.** The agent submits an action request. A human reviews and approves. No exceptions for speed.
- **Break-glass monitoring.** If an agent attempts an action outside its normal pattern, alert a human immediately. Not after the data is gone.
- **Audit trails that cannot be modified by the agent.** The agent should not have write access to its own logs. If you are relying on the agent to self-report its actions, you have already lost.
What about detection?
Current detection capabilities for agent-based attacks are primitive. Most injected prompts do not match known signatures. Behavioral baselines are the best bet. Monitor tool call frequency, target endpoints, data volumes, and timing patterns. If your agent has never made a database query at 3 AM and suddenly it is dumping the customer table, something is wrong.
But baseline detection requires learning what normal looks like, and most agent deployments are too new to have reliable baselines. You are flying blind until you have collected enough telemetry.
The bottom line
The agent supply chain is not a future problem. It is a present threat. Indirect prompt injection works today. Tool access escalation works today. Supply chain poisoning through model repos and plugins works today. These are not theoretical attack patterns. They have been demonstrated, documented, and in some cases, exploited in production.
Every integration you add to your agent is an attack surface. Every plugin is a potential backdoor. Every API key you give the agent is a credential that can be exfiltrated through a compromised chain-of-thought trace.
The security community spent a decade learning that software supply chains needed hardening. The AI industry is about to learn the same lesson in months. The teams that treat agent security as a first-class engineering requirement, not an afterthought handled by a system prompt, will be the ones that do not get pwned.
Everyone else will learn the hard way.
Victor Kane is a Security and OSINT Analyst. This analysis is based on published research, disclosed vulnerabilities, and observed attack patterns as of June 2026.