
Short answer: Agentic AI security protects AI systems that act on their own by calling tools, reading data and writing code. The core threats are prompt injection, over-permissioned tools, leaked credentials and data exfiltration. The core controls are scoped credentials per agent, human approval for high-impact actions, allowlisted tools and egress, sandboxing and a log of every tool call.
Related guides:
- Millions of AI agents are running without oversight. Is yours one of them?
- Your auditor is about to ask about AI agents: 9 things they'll want to see
- How to Strengthen Your AI Security Posture Before Hackers Exploit It
Key takeaways
- An AI agent is a new kind of insider: it holds credentials, touches production data and acts faster than any human reviewer can watch.
- Any content an agent reads, including emails, web pages, tickets and documents, can carry instructions that hijack it. Treat all of it as untrusted input.
- The blast radius of an agent is set by its permissions, not by its prompt. Least privilege is the control that matters most.
- Log every tool call with the agent identity, the input, the output and the human who approved it. That log is both your detection layer and your audit evidence.
- Many agent risks map onto SOC 2, ISO 27001 and ISO 42001 controls you already run, once you extend them to non-human identities.
Why is agentic AI security different from chatbot security?
Agents are different because they do things, while a chatbot only says things. A chatbot that gets tricked produces a bad answer. An agent that gets tricked sends the email, runs the query, merges the pull request or issues the refund.
Four properties change the threat model:
- Autonomy. Agents decide which step to take next without a human reviewing each one. Mistakes compound before anyone notices.
- Tool access. Agents call APIs, databases, file systems, browsers and shells. Every tool is a capability an attacker can borrow.
- Chained actions. One task can involve dozens of tool calls. A single poisoned input early in the chain can steer everything that follows.
- Memory. Many agents persist context across sessions through vector stores, notes files or conversation history. Whatever lands in memory can influence future runs, including runs for other users.
Governance answers who approved an agent and who owns it; our guide on governing AI agents covers that. This post covers what can go wrong technically once the agent is running.
What are the biggest AI agent security risks?
The biggest risks come from untrusted input reaching an agent that holds more power than its task requires. The OWASP Top 10 for LLM Applications lists prompt injection and excessive agency among its core risks, and both show up in almost every agent incident pattern. Here is the fuller threat model.
Direct and indirect prompt injection
Direct injection is a user typing instructions that override the agent's rules. Indirect injection is worse: the instruction hides inside content the agent reads, such as a ticket, PDF, web page or code comment. The agent cannot reliably tell data from instructions.
Excessive agency and over-permissioned tools
Teams often give an agent a broad API key "to get it working" and never narrow it. An agent that only reads invoices but can write to billing turns every injection into a financial event.
Credential and secret exposure
Secrets that agents use to call tools end up in prompts, logs, memory stores and model outputs. Coding agents can commit them to repositories or print them into build logs.
Data exfiltration through tool outputs
Attackers only need the agent to send data out, for example by rendering a URL with data in the query string, posting to a webhook or emailing an external address.
Agents as a new kind of insider
An agent is a non-human identity with standing access, running around the clock. If nobody owns it, nobody reviews its access, rotates its credentials or notices drift.
Memory and context poisoning
Anyone who can write to an agent's long-term memory or retrieval corpus can plant instructions or false facts that surface weeks later, while outputs look normal.
Supply chain risk from plugins and tool servers
Third-party plugins and MCP-style tool servers are code you did not write, sitting directly in the agent's decision loop. Even a tool description can carry injected instructions.
Runaway cost and loops
Agents can get stuck retrying or calling paid APIs in a loop, an availability and financial risk that attackers can trigger deliberately.
How do you secure AI agents? Ten technical controls
You secure AI agents by limiting what they can do, checking what goes in and out, and recording everything they touch. No single control is enough, because prompt injection has no complete fix today. Design so that a fully hijacked agent still cannot do serious damage.
| # | Control | What it looks like in practice |
|---|---|---|
| 1 | Least-privilege, scoped credentials per agent | One identity per agent, short-lived tokens, read-only by default |
| 2 | Agents as non-human identities with owners | Named human owner, included in access reviews and offboarding |
| 3 | Human approval for high-impact actions | Payments, deletions, external emails, deploys and bulk exports need a person to confirm |
| 4 | Allowlisted tools | Only registered, reviewed tools with pinned versions |
| 5 | Allowlisted egress | Egress limited to known domains, so exfiltration fails |
| 6 | Input and output filtering | Scan inputs for injection, outputs for secrets and sensitive data |
| 7 | Sandboxing | Code execution and browsing in isolated containers without production credentials |
| 8 | Log every tool call | Agent ID, user, tool, parameters, result, approver and timestamp |
| 9 | Red-teaming | Planting indirect injection payloads in the sources agents read |
| 10 | Kill switch and budgets | Disable an agent and revoke its credentials in minutes; rate limits and spend caps |
Two design rules make the rest work. First, let the model propose actions but have deterministic code validate them against policy before execution. Second, inject credentials at the tool layer so the model never sees raw tokens.
Mapping agent threats to controls and audit evidence
Most agent threats map onto controls you already operate for SOC 2, ISO 27001 or ISO 42001, so the job is to extend them rather than invent new ones. The table below shows which control addresses each threat and what evidence an auditor can inspect. Framework references are described generally. Confirm the exact criteria and control numbers with your auditor.
| Threat | Primary control | Evidence to retain | Framework area |
|---|---|---|---|
| Prompt injection (direct and indirect) | Input filtering, output validation, human approval gates | Filter configuration, red-team reports, approval records | SOC 2 system operations and monitoring; ISO 27001 secure development; ISO 42001 AI risk treatment |
| Excessive agency | Least-privilege scoped credentials | IAM policies per agent, access review sign-offs | SOC 2 logical access (CC6); ISO 27001 access control and access rights; ISO 42001 AI system design controls |
| Credential and secret exposure | Secrets manager, tool-layer injection, secret scanning | Rotation logs, scanner results, vault access logs | SOC 2 logical access; ISO 27001 authentication information and cryptography |
| Data exfiltration | Egress allowlist, DLP on outputs | Network policy, DLP alerts and dispositions | SOC 2 confidentiality criteria; ISO 27001 data leakage prevention; ISO 42001 data for AI systems |
| Agent as unowned insider | Non-human identity inventory with owners | Identity inventory, owner assignments, offboarding records | SOC 2 logical access; ISO 27001 identity management; ISO 42001 roles and responsibilities |
| Memory and context poisoning | Write controls and review on memory and retrieval sources | Access policies on vector stores, content review logs | ISO 27001 information integrity controls; ISO 42001 data quality and provenance |
| Plugin and tool server supply chain | Tool allowlist, vendor risk review, version pinning | Vendor assessments, approved tool register, change tickets | SOC 2 vendor management and change management (CC8); ISO 27001 supplier relationships; ISO 42001 third-party relationships |
| Runaway loops and undetected misuse | Tool call logging, spend caps, kill switch | Tool call logs, budget alerts, kill switch tests | SOC 2 monitoring (CC7); ISO 27001 logging and monitoring; ISO 42001 operational monitoring |
For a closer look at how auditors test these areas, see what auditors want to see on AI agents.
HealthTech example: an agent with access to PHI
A patient-intake agent shows how these controls stack up. Picture a HealthTech company whose agent reads referral emails, extracts patient details, checks insurance eligibility and drafts appointments.
The attack. A referral PDF hides text telling the agent to look up the provider's last 50 patients and email their names and dates of birth to the sender. An agent with broad EHR access and unrestricted email could comply, causing a breach of protected health information.
The layered defense:
- Scoped access. The agent can read only the record linked to the current referral, in line with the HIPAA minimum necessary standard.
- Egress and recipient allowlists. Email goes only to verified provider domains, and output filtering blocks bulk PHI.
- Human approval. Any message leaving the organization that contains PHI goes to a staff member for one-click review.
- Sandboxed parsing. Documents are converted to plain text in an isolated service, and embedded instructions are flagged before the content reaches the model.
- Audit trail. Every record opened and message drafted is logged with the agent ID, supporting HIPAA Security Rule audit controls.
- Business associate coverage. If the model provider or any tool server processes PHI, a business associate agreement needs to be in place before go-live.
The injection may still reach the model. It just cannot turn into a breach, because the agent never held the access or the channel to carry it out.
Where should a small team start?
Start by finding every agent you already run and cutting its permissions to what its task needs. A first-month checklist:
- Inventory every agent and automation that calls a model and takes actions.
- Assign each a human owner in your identity inventory.
- Replace shared API keys with a scoped credential per agent.
- Remove tools each agent does not need.
- Put an approval step in front of high-impact actions.
- Send tool call logs to your security logging stack.
- Add plugins and tool servers to your vendor register.
- Test the kill switch in a tabletop exercise.
- Red-team indirect prompt injection through each agent's data sources.
For broader hardening beyond agents, our list of AI security practices every business needs covers model, data and vendor controls.
How SecureSlate helps
SecureSlate helps you treat AI agents as part of your existing security program instead of a separate project. You can add agent risks to your risk register, assign owners, include agent service accounts in access reviews, and map agent controls such as scoped access, approval gates and tool call logging to a single control library that you track across SOC 2, ISO 27001 and ISO 42001. Automated evidence collection from your cloud and SaaS stack keeps access and logging evidence current, and your AI acceptable use and agent access policies live in the same policy library as the rest of your program.
Plugins, model providers and tool servers can go through vendor risk management, and audit-ready exports plus a trust center help you show customers how your agents are controlled.
Start your free SecureSlate trial
FAQ
What is agentic AI security?
It is the set of threats and controls for AI systems that take actions, such as calling APIs, reading internal data or writing code. It focuses on limiting what agents can do, validating inputs and outputs, and recording actions.
Can prompt injection be fully prevented?
Not with current techniques. Filters and model-level defenses reduce the risk, but a determined attacker can often find a phrasing that gets through. That is why the most reliable controls limit the damage a hijacked agent can cause, through least privilege, egress limits and human approval.
Do SOC 2 and ISO 27001 cover AI agents?
Neither framework names AI agents specifically, but their access control, change management, vendor management and monitoring requirements apply to agents like any other system component. ISO 42001 adds AI-specific risk and impact assessment on top. Expect auditors to ask how agents appear in your inventory and access reviews.
Disclaimer (legal note)
This article is for general information only and is not legal, regulatory or professional advice. Requirements vary by framework, industry and jurisdiction. Consult qualified advisors for your specific obligations.
Need compliance without the complexity?
SecureSlate automates ISO 27001, SOC 2, GDPR, HIPAA, and more. Built for growing teams. See it in action.
Find compliance gaps in 30 seconds
