Prompt injection
Prompt injection is an attack in which instructions hidden in content an AI system processes, such as a web page, email, document, or user message, trick the model into ignoring its original instructions. It can make an AI leak data, take unintended actions, or produce harmful output. OWASP lists it first among security risks for LLM applications.
01why it matters for a business
The risk grows with what the AI can do. A chatbot that can only answer questions might be tricked into saying something embarrassing. An agent that reads incoming email and can send messages, query a database, or browse the web might be tricked into forwarding confidential data to an attacker, because the attacker's instructions arrived inside an email the agent was asked to process. Models cannot reliably tell instructions from you apart from instructions embedded in data.
There is no complete fix at the model level today, so defense is architectural. Limit what any single agent can access and do. Treat everything the model reads from outside as untrusted. Require confirmation for sensitive actions. Separate agents that read untrusted content from agents that hold privileged access. Watch for unusual patterns in tool calls.
02what it looks like in practice
A recruiting team uses an AI agent to screen resumes and draft summaries for hiring managers. A candidate hides white-on-white text in a resume telling any AI reader to rate the candidate as an exceptional fit. A well-designed system limits the damage: the agent only produces a summary scored against the job criteria, flags any text addressed to an AI, and a recruiter reviews every shortlist. A poorly designed agent with write access to the applicant tracking system and email could be pushed into far more.
03common mistakes
- Believing a stronger system prompt solves it. Instructions can be overridden by other instructions.
- Giving an agent that reads untrusted content the ability to send data outside the company.
- Connecting third-party tools and data sources to agents without reviewing what content flows through them.
- No logging, so an injection goes unnoticed.
04related terms
- AI guardrailsGuardrails are the controls around an AI system that keep its inputs, outputs, and actions within acceptable limits: input filtering, output checks for policy, format, and sensitive data, limits on which tools and data it can reach, spending caps, and rules that route risky cases to a person.
- System promptA system prompt is the standing set of instructions an application gives a language model before any user input, defining its role, rules, tone, available tools, and limits.
- Model Context Protocol (MCP)The Model Context Protocol (MCP) is an open standard, introduced by Anthropic in 2024, for connecting AI applications to external tools and data.
- AI agentAn AI agent is software that uses a language model to pursue a goal by choosing its own next steps: it reads the situation, picks a tool or action, checks the result, and repeats until the task is done or it needs a person.
- Tool use (function calling)Tool use, also called function calling, is the ability of a language model to request that your software run a specific function, such as looking up an order, querying a database, or sending a message, with structured inputs the model fills in.
05where insomnia club fits
Insomnia Club designs agents with prompt injection in mind from the first architecture sketch: least-privilege tools, untrusted-content handling, confirmation steps, and logging.
see AI agent development →tell us what keeps you up at night.
Scoped by the people who ship it. Priced before we start.
book a call drop your number