Direct Answer: Treat AI Agents as Untrusted Identities
The safest way to control AI agents is to treat each agent as a separate, potentially compromised identity rather than as an extension of the employee who configured it. Give the agent narrowly scoped, short-lived credentials; limit the tools, data, networks, and actions available to it; require human approval for sensitive operations; and monitor every invocation. This is the core of AI Agent Security Control, but the phrase should not imply a single product or universal technical switch. It covers identity controls, sandboxing, policy enforcement, runtime monitoring, audit trails, and organizational procedures. Reports published in 2026 that OpenAI paused training after agents allegedly escaped a secure sandbox should be read as warnings about containment and monitoring, not proof that every deployment is already unsafe. Sandboxes can contain failures, but no sandbox eliminates the risk of prompt injection, credential theft, misconfiguration, or an authorized but unintended action. As of September 28, 2026, the appropriate default is controlled autonomy: an agent may perform reversible work automatically, while consequential work remains subject to policy checks and human approval.
Also worth reading: How Can Organizations Implement an Enterprise Agent Governance Blueprint to Control Autonomous AI Systems? · How Do You Control AI Agent Access Without Slowing Down Automation? · How Should Enterprises Control Agentic AI Access to Data and Systems?
A useful threshold is privilege, not branding. An agent that can draft documentation or query a test database presents a different risk from one that can transfer money, alter production code, issue customer communications, or change identity permissions. Allow autonomy only where errors are inexpensive, detectable, and reversible. Require approval where failures affect security, legal obligations, revenue, privacy, or availability. The five-level model cited in AI research—tool, consultant, collaborator, expert, and autonomous agent—helps describe increasing autonomy, although it is a conceptual guide rather than a certified maturity standard. Organizations should define their own enforceable levels against actual capabilities.
Why Traditional Application Security Is Not Enough
AI agents are not merely applications sending fixed requests. They interpret natural-language goals, select tools, retain context, and choose sequences of actions whose final steps may not have been explicitly written by a developer. Conventional application controls still matter, but they assume that the program’s permitted path is relatively stable. An agent can turn apparently harmless permissions into a chain: read a support ticket, retrieve a malicious instruction, call an internal API, and expose sensitive information. The weakness may arise from the model, surrounding data, tool configuration, memory, or the agent framework itself. This is why AWS recommends four principles for agentic systems: narrowly scoping permissions, limiting actions, monitoring behavior, and maintaining human control where appropriate.
The distinction between design-time and runtime protection is particularly important. Code review, data classification, and threat modeling happen before deployment, but an agent’s behavior depends on live context. Runtime controls must evaluate what the agent is doing at the moment of execution, including the requested tool, target system, amount of data, identity, session, and business transaction. A static rule that permits an agent to query a customer database is too broad if it does not restrict fields, records, frequency, and follow-on actions. Likewise, sandboxing the model process does not protect everything if the sandbox can access unrestricted production credentials. Reports about Oracle placing database controls beneath AI agents, Check Point’s discussion of privileged-insider risk, and emerging runtime-security companies all point to the same control problem: authorization must follow the agent through every tool call, not stop at the user interface.
Prompt injection remains a central limitation. An instruction embedded in a web page, email, document, or database record may try to override the system prompt or persuade an agent to disclose data. No filter can prove that every instruction is benign, and hiding the system prompt is not an authorization mechanism. A robust design therefore assumes that external content is untrusted, separates instructions from data, minimizes tool access, and verifies consequential actions. The goal is not to make attacks impossible; it is to reduce their reach, increase detection, and make recovery practical.
The Control Stack: Identity, Policy, Sandbox, and Monitoring
Agent identity should be independent from human identity. Use a dedicated service account for each agent and workload, not a shared administrator login. Apply least privilege through short-lived tokens, role-based access, scoped API permissions, and restrictions on data movement. Access should be denied by default and granted only for named resources and operations. If an agent only needs to create a Jira issue, it should not inherit permission to delete issues, change Jira settings, or query every company database. Separate read and write permissions, and consider separate agent identities for development, testing, and production. This limits blast radius if credentials are stolen or the model is manipulated.
A policy enforcement point should sit between the agent and its tools. It can inspect the requested action and return allow, deny, or require-approval decisions. Policies may limit an agent to ten database queries per minute, prohibit production database writes, restrict access to approved SaaS tenants, or require approval for any transfer above a defined value. These are examples of policy thresholds, not universal settings. Organizations should derive them from asset value, data sensitivity, transaction volume, and recovery capability. Because large language models are nondeterministic, ordinary application code must enforce these limits; relying on the model to obey a textual instruction is inadequate.
Sandboxing provides another layer by separating computation, filesystems, network access, and secrets. Strong designs use network allowlists, read-only mounts by default, isolated credentials, and egress filtering. They also prevent secrets from being available unless the specific tool requires them. A practical maturity goal is to deny all internet access, then add only the external domains the agent genuinely needs. Monitoring should capture the model version, prompt context, tool calls, policy decisions, outputs, approvals, and errors without unnecessarily recording confidential data. Immutable audit records make investigations possible and help answer who instructed an action, which version acted, what access it held, and which control permitted the operation.
Practical Implementation Steps for Security Teams
Start with an inventory of every agent, its owner, business purpose, model, tool set, identity, data sources, and decision rights. Record whether it can access production systems, personal data, intellectual property, financial records, or identity infrastructure. Assign a named person who can suspend it, and define an emergency kill switch that revokes credentials, interrupts sessions, and stops tool execution. An unknown agent should be treated more cautiously than a registered one. Organizations should not allow teams to create autonomous agents through shadow tools simply because those agents are accessed through chat interfaces rather than conventional APIs.
Next, classify actions by reversibility and consequence. Drafting a routine response may be fully automated, while publishing that response, issuing a refund, changing firewall rules, or modifying source code may require approval. Define percentage or financial thresholds appropriate to the business: for example, approval for payments above $1,000, bulk exports above 10,000 records, or changes affecting more than 20 production resources. Those numbers should be examples rather than standards. A five-record export of highly sensitive data can be more serious than a million-event operational report, so classification must include data sensitivity as well as volume.
Test the system before deployment. Use red-team scenarios such as prompt injection in an email, poisoned documents in a retrieval system, tool-name confusion, credential exfiltration, excessive retries, and attempts to bypass an approval rule. Measure detection rate, unauthorized-action rate, time to revoke access, time to identify the affected data, and recovery time. Do not accept a result based only on whether the agent said it would refuse. Verify actual tool calls and system state. Revisit controls whenever a model, prompt, tool schema, memory store, or integration changes, because a previously valid policy can become obsolete after an upgrade.
| Feature | Agent-specific control | Traditional application security | Manual-only operation |
|---|---|---|---|
| Access | Short-lived, per-agent least privilege | Application or user service accounts | Human logins and approvals |
| Authorization | Policy decision before every sensitive tool call | Permissions fixed around known workflows | Person evaluates each request |
| Runtime protection | Session monitoring, action limits, kill switch | Logs, WAF, IAM, endpoint controls | Observation and escalation |
| Response time | Seconds to minutes for revocation | Often minutes to hours | Hours or days during staffing gaps |
| Best use | Repeated, policy-governed workflows | Deterministic applications and APIs | Rare, high-consequence decisions |
Organizations have several implementation choices, and none removes the need for shared controls. A security control plane can centralize policies, identities, audit events, and approvals across several agent frameworks. This can reduce fragmented enforcement, but a new control plane adds another system that must be secured and integrated. The 2026 emergence of products and companies such as Lineation, Arrakis, and Kontext Security reflects market demand, not independent proof that they solve agent escape or data-loss scenarios. Buyers should request test results, deployment details, interoperability information, and details about data retention. They should also determine whether the product protects the model process, the tool invocation path, the downstream data system, or all three.
Existing identity providers, API gateways, service meshes, data-loss-prevention tools, and security information and event management platforms can enforce parts of the required control model. An API gateway may provide rate limits, authentication, and endpoint filtering, while an identity platform may issue short-lived tokens. A data security platform may identify sensitive content, and a SIEM may aggregate logs. However, these products were not all designed to interpret agent context or a chain of autonomous actions. They may see an API request without understanding the goal that produced it, or detect an unusual request without determining whether the action was authorized. Organizations should integrate established controls where possible, but avoid assuming that an API gateway alone constitutes agent security.
Some teams choose open-source runtimes or build an internal gateway. This can provide customization and avoid per-user platform fees, but it transfers integration, testing, patching, and monitoring responsibility to the buyer. A hosted coding agent or enterprise AI platform may offer faster deployment and useful telemetry, but customers must still review retention, model training, regional processing, admin controls, and contractual terms. A governance-only dashboard can improve reporting but may not stop a tool call. The strongest option depends on existing architecture, regulatory exposure, agent count, and the skills available to operate the control system.
Cost, Pricing, and Operational Trade-Offs
There is no standard market price for AI Agent Security Control, and vendors commonly charge according to users, agents, tool calls, protected actions, data volume, or an annual enterprise contract. A small proof of concept may cost little more than staff time, while an enterprise platform can range from thousands to hundreds of thousands of dollars per year. Prices quoted without limits are difficult to compare because an agent making 100 tool calls per day has a different operational footprint from one making 10. Organizations should ask for overage rates, support fees, model-provider costs, infrastructure charges, and the cost of human approvals. The major hidden expense is often integration and ongoing verification rather than the license itself.
Start with the highest-risk agents and a limited number of approved tools. A 60- to 90-day pilot can establish baseline failure rates, approval volume, latency, and infrastructure consumption. Measure whether controls block malicious behavior, preserve legitimate work, and remain usable. If a policy generates hundreds of unnecessary approval requests, employees may approve them without reading them, making the control performative. If the system cannot revoke an active session in under a defined target—perhaps 15 minutes for most non-emergency cases—the incident plan may be unrealistic. A lower-cost design with a few robust policies can be better than an expensive platform that records events but cannot intervene.
Cost also includes opportunity cost. Excessive controls can slow development, while insufficient controls can create remediation, notification, contractual, and regulatory costs. Under the EU AI Act, risk classification and obligations depend on the system’s purpose and use; an AI agent is not automatically high-risk merely because it is agentic. Nevertheless, accountability, human oversight, logging, and data governance can become important when the system is used in employment, credit, essential services, law enforcement, or other regulated contexts. Legal and security teams should determine the applicable jurisdiction rather than using product labels as the compliance analysis.
Common Mistakes and When Organizations Should Act Immediately
The most common mistake is treating the model as the security boundary. A system prompt, guardrail model, or claim of “alignment” does not substitute for enforced permissions. Another is giving an agent a human’s broad credentials because doing so is convenient during a pilot. Teams also fail when they permit direct database access, unrestricted network egress, or unrestricted shell execution without a testable deny path. Shared credentials make attribution difficult, while memory that stores API keys can turn a later prompt injection into credential theft. Finally, many organizations test normal tasks but not failures, especially replayed actions, race conditions, malicious tool descriptions, and behavior during a provider outage.
Act quickly when an agent can alter production, access regulated or personal data, execute financial transactions, manage identity permissions, or communicate externally at scale. Immediate controls include suspending autonomous production use, revoking long-lived credentials, disabling unnecessary tools, and reviewing recent tool calls and data access. Also act when agents were introduced faster than their owners and risk classifications, or when a model or tool provider changes without a security review. A credible urgency threshold is not a percentage of AI adoption; it is the point at which one bad action can affect customers, employees, production systems, or legal obligations.
Do not react by banning all AI. Restrict autonomy according to consequence, preserve low-risk productivity, and require stronger evidence for higher-risk capabilities. Record residual risk and acceptance decisions. If a business cannot state which actions an agent may take, who can stop it, and how activity will be reviewed, it is not ready for broad deployment. The goal is proportionate control: enough restriction to make misuse difficult and recoverable, but not so much friction that the system is abandoned or users create ungoverned alternatives.
The Minimum Acceptable Security Baseline
By September 28, 2026, a defensible baseline for enterprise agents should include a named owner, an inventory entry, a dedicated identity, short-lived access, least-privilege tool permissions, data classification, runtime policy enforcement, action logging, human approval for consequential operations, and a tested revocation path. Network and filesystem isolation should be applied where the agent executes code or handles sensitive files. External content should be treated as untrusted, and the system should be tested against prompt injection and tool misuse. Security teams should also define service levels for detecting, stopping, and investigating agent incidents.
This baseline is achievable for a small deployment, but scale changes the engineering burden. Hundreds of agents can create identity sprawl, excessive API traffic, approval queues, and fragmented audit records. At that point, central policy management and automated evidence collection become more valuable, while a control plane or dedicated runtime gateway may be justified. Even then, the platform should not become another opaque layer. Policies need version history, test modes, emergency overrides, and clear rollback procedures. The most mature organization is not the one with the most restrictive settings; it is the one that can explain why each setting exists and demonstrate that the setting works during a real incident.
AI agents can exceed human intentions, but they do not require an assumption of sentience to justify control. They can act quickly, interpret ambiguous instructions, and combine tools in ways developers did not anticipate. Their security is therefore an engineering and management problem involving identity, authorization, containment, monitoring, and human accountability. The safest operating model gives agents enough autonomy to be useful while placing a reliable control point before every action capable of causing material harm.