Glossary Background Image

No Bad Questions About ML

Definition of Prompt injection

What is prompt injection? Types, risks, and examples

Prompt injection is a security vulnerability in LLM-based applications where untrusted input is interpreted as instructions and changes the model’s intended behavior. Depending on the surrounding system, this can lead to unwanted outputs, data exposure, policy bypass, or unintended actions through connected tools.

The risk comes from the way LLM applications process instructions and ordinary content within the same natural-language context. Prompt injection becomes more consequential when a model can access private information, retrieve external content, call APIs, or perform actions on a user's behalf.

What is prompt injection?

Prompt injection occurs when content an AI system should treat as data instead influences what the model does. The injected instruction may come directly from a user or indirectly through content the application retrieves or processes.

The surrounding application determines how serious the effect can become. In a simple chatbot, an injection may only change an answer. In a RAG system or AI agent with access to internal data and external tools, manipulated behavior can have broader consequences because the model may retrieve information, call services, or initiate actions.

Prompt injection can affect more than model output. The permissions, data access, integrations, and controls surrounding the LLM determine how much impact a successful injection can have.

How does prompt injection work?

Prompt injection exploits a trust problem inside the context provided to an LLM.

An application may combine trusted instructions from developers with user input, retrieved documents, webpages, emails, or tool results. From the model’s perspective, much of this information arrives as text or other model-readable content.

If untrusted content contains instruction-like language and the model treats it as authoritative, its behavior can deviate from the original task. It may follow the new instruction instead of, or alongside, the application’s intended rules.

This differs from traditional SQL or command injection. Those attacks exploit how an interpreter handles executable syntax, while prompt injection manipulates how an LLM interprets natural-language context.

Direct vs indirect prompt injection

The main distinction between direct and indirect prompt injection is where the manipulative instruction enters the AI system.

Direct prompt injection

Direct prompt injection comes through input controlled directly by the user interacting with the model. The input attempts to alter the model’s intended behavior, such as making it disregard application instructions, change its task, or disclose context it should not reveal.

The familiar "ignore previous instructions" pattern falls into this category, but direct injection isn't limited to one phrase or technique.

Indirect prompt injection

Indirect prompt injection enters through external content that the AI system consumes while completing another task.

For example, an AI application may read:

  • Webpages
  • Documents
  • Emails
  • Retrieved RAG passages
  • Database content
  • Tool outputs

If one of those sources contains instruction-like content, the model may treat it as part of the task rather than as untrusted data.

RAG applications, browsing systems, document-processing assistants, email tools, and AI agents are especially exposed to this pattern because they routinely consume content that neither the user nor the application developer fully controls.

Direct injection comes through user interaction; indirect injection comes through content the system processes while performing the task.

What are the risks of prompt injection?

Prompt injection does not automatically compromise an entire system. Its impact depends on what information, permissions, integrations, and actions are available to the affected AI application.

Sensitive data exposure

A manipulated model may disclose information included in its context or retrieved from connected systems. Depending on the application, this could include internal documents, customer information, proprietary data, or other material the model is authorized to access for legitimate tasks.

Instruction and policy bypass

Prompt injection may cause the model to disregard application-level rules or operate outside the workflow developers intended. This can weaken business restrictions, content rules, or safeguards that depend primarily on the model following instructions correctly.

Unauthorized actions

The potential impact grows when an AI agent can call tools or APIs. Manipulated behavior may influence actions such as retrieving records, sending information, updating data, or initiating workflows.

Prompt injection does not automatically grant an attacker arbitrary system access. The consequences are limited or amplified by the permissions and authorization controls around the model.

Manipulated or unreliable output

An injection can also distort summaries, recommendations, classifications, or other outputs without accessing sensitive information.

For example, attacker-controlled content retrieved by an AI assistant could influence how the system summarizes a document or evaluates information. If downstream users or automated workflows trust that output without further checks, the manipulation can affect business decisions or process results.

Prompt injection vs jailbreaking

Prompt injection and jailbreaking overlap, but they describe different security concerns.

Prompt injection concerns untrusted instructions changing the intended behavior of an LLM-based application. Jailbreaking usually refers more specifically to attempts to bypass model restrictions or safety behavior.

A jailbreak can be part of a prompt injection scenario, but prompt injection also includes cases such as indirect instructions in retrieved content that redirect an application without targeting model-level safety controls.

Prompt injection also differs from data poisoning: prompt injection affects behavior through content processed at inference or runtime, while data poisoning targets data used to train or otherwise modify model behavior.

Example: indirect prompt injection in a RAG assistant

Consider an internal AI assistant that retrieves company documents and summarizes them for employees. One retrieved document contains instruction-like text that conflicts with the user’s original request.

A weakly controlled application may allow that content to influence the model’s behavior instead of treating it only as source material. A more controlled design treats retrieved content as untrusted, restricts what information and tools the model can access, and applies separate controls to sensitive actions.

Adding RAG does not by itself remove prompt injection risk.

Can prompt injection be prevented?

Prompt injection risk can be reduced, but a system prompt, input filter, or model refusal should not be treated as a complete security boundary.

A defense-in-depth approach typically includes:

  • Treat external content as untrusted. User input, retrieved documents, webpages, emails, and tool results should not receive the same trust as application instructions.
  • Limit permissions. Models and agents should receive only the data and tool access needed for their tasks.
  • Enforce controls outside the model. Sensitive operations should rely on application-level authorization and validation.
  • Add approval for high-impact actions. Human confirmation can reduce risk before sensitive or difficult-to-reverse operations.
  • Test direct and indirect injection paths. Evaluation should cover user input, retrieval, browsing, documents, and integrations.
  • Monitor failures and unusual behavior. Production systems should support investigation when suspicious interactions occur.

Prompt injection should also be considered during threat modeling because its impact depends on trust boundaries, available data, permissions, and possible downstream actions.

Inline CTA icon

Prompt injection risk depends on the whole AI system

Retrieval, tool permissions, access controls, integrations, and approval logic can determine whether manipulated model behavior remains an incorrect response or develops into a broader security issue.

Explore AI/ML technical audit →

Key Takeaways

  • Prompt injection occurs when an LLM-based application interprets untrusted content as instructions and changes its intended behavior.
  • Direct prompt injection enters through user-controlled input, while indirect injection reaches the model through documents, webpages, RAG content, emails, or tool outputs.
  • The potential impact grows when an AI system can access sensitive information, call external tools, or perform actions in connected systems.
  • Prompt injection can cause data exposure, policy bypass, unintended actions, or manipulated outputs, but its consequences depend on surrounding permissions and architecture.
  • Prompt injection and jailbreaking overlap, but jailbreaking more specifically targets model restrictions, while prompt injection can manipulate broader application behavior.
  • No single system prompt or filter provides complete protection, so reducing risk requires layered controls around data, permissions, actions, testing, and monitoring.

FAQ

What is a prompt injection attack?

A prompt injection attack attempts to make an LLM-based application treat untrusted input as instructions and behave differently from what its developers or users intended. The instruction can be provided directly or embedded in content the system later processes.

Can prompt injection affect RAG systems?

Yes. Retrieved content can contain instruction-like text that influences the model after it is added to the context. RAG can improve access to relevant information, but it does not by itself eliminate prompt injection risk.

Is prompt injection the same as jailbreaking?

No. The concepts can overlap, but jailbreaking generally targets restrictions or safety behavior of the model itself. Prompt injection is broader and can redirect an application without necessarily attempting to bypass model-level safety controls.

Can prompt injection be completely prevented?

There is no single control that reliably eliminates the risk across all LLM applications. Organizations can reduce potential impact by limiting permissions, treating external content as untrusted, enforcing authorization outside the model, testing injection scenarios, and requiring approval for sensitive actions.