An AI chatbot can give a bad answer. An AI agent can give a bad answer and then act on it. That difference is why agent security is becoming one of the most urgent problems in software development.
Imagine a customer-support agent that can read an order, calculate a refund, update a database, and send confirmation. The automation is useful because it removes repetitive work. It is also dangerous if a malicious customer convinces the model to refund far more than the purchase price or to run code that exposes environment variables.
A system prompt saying "never issue an excessive refund" is not a sufficient defense. Prompts are instructions interpreted by a probabilistic model. They can be overridden by prompt injection, weakened by tuning, misapplied in an unusual context, or behave differently after a model update. Money, secrets, and production data need enforcement outside the model.
Google's proposed pattern is a zero-trust architecture for AI agents. The model is allowed to reason flexibly, but the surrounding infrastructure assumes it can be tricked. Every sensitive action must prove its identity, untrusted code runs inside strong isolation, and deterministic gateways block requests that violate business rules.
This article explains the three layers in Google's design, why each one is necessary, and how teams can adapt the idea even when they are not using Google Cloud or the Agent Development Kit.
The Key Takeaways
- An autonomous agent connected to databases and tools is part of the production system, not merely a text generator.
- System prompts are useful behavioral guidance but should not be treated as hard security boundaries.
- Google's pattern combines cryptographic signatures for state-changing actions, gVisor sandboxing for generated code, and deterministic semantic gateways for inputs and outputs.
- Each layer solves a different failure: attribution and tamper evidence, runtime containment, and business-policy enforcement.
- Least privilege, human approval, logging, testing, and incident response remain necessary around the three core controls.
- The zero-trust principle applies to any agent framework, not only Gemini or Google's ADK.
From Chatbot to Production Actor
A chatbot generally receives text and returns text. It may be wrong, but a human decides what to do next. An agent can receive a goal, select tools, inspect data, call APIs, generate code, update records, and continue until it believes the goal is complete.
This is the shift described in our guide to what AI agents can actually do in 2026. The agent's value comes from reducing human coordination. It also receives authority that a traditional assistant did not have.
Once an agent can mutate production state, its output is no longer only a message. A database write, refund, deleted file, deployed patch, sent email, or changed permission can have an immediate and sometimes irreversible effect. Security teams must treat the model, tools, session state, credentials, and generated code as one application.
The risk is not limited to a malicious model. A well-intentioned model can misunderstand a user, follow hidden instructions in retrieved content, call the wrong tool, reuse stale context, or produce vulnerable code. Zero trust assumes failure is possible and limits its impact.
Google's Refund-Agent Scenario
Google's reference design uses an autonomous customer-support and returns agent built with Gemini and the Agent Development Kit. In normal operation, the agent reads a request, calculates a prorated deduction, writes an approved refund to a ledger, and returns a receipt.
An attacker can hide a prompt injection inside the customer message. The message may tell the agent to ignore prior instructions, increase the refund far beyond the order value, sign the transaction, and run a script that prints environment variables. If the agent uses a shared database connection and runs code in an ordinary container, one successful injection could create an unauthorized payment and expose API keys.
The scenario is deliberately simple because the same pattern appears in higher-stakes systems. A procurement agent could approve a false vendor. A coding agent could send repository secrets to an external server. A calendar agent could invite the wrong people. A finance agent could alter a payment destination. A support agent could reveal another customer's records.
The security question is not "Did we tell the agent to be careful?" It is "What technically prevents the agent from completing an unauthorized action even when its reasoning is wrong?"
Why System Prompts Are Not Security Boundaries
A system prompt is part of the model's context. It can define a role, priorities, format, refusal behavior, and workflow. It is valuable, but it is still language interpreted by the same model that interprets untrusted user input.
Prompt injection exploits that shared interpretation layer. A malicious instruction may arrive directly from a user, indirectly through a webpage, inside a document, in a database field, or from another agent. The model must distinguish trusted policy from untrusted content, but complex tasks can blur the boundary.
Even without an attacker, prompts can drift. Developers edit them. Models change. Routing sends a request to a different model. Long context pushes relevant instructions far away. A rare situation may conflict with the examples used during testing.
Security controls should therefore be deterministic whenever the rule is deterministic. "Never refund more than the verified order amount" is arithmetic and authorization logic. Ordinary software can enforce it exactly. The model can propose a refund, but the transaction service should reject an amount outside the allowed range.
Google's ADK safety guidance makes a similar point: tools can expose only the actions a model is allowed to take, and in-tool guardrails can deterministically eliminate entire classes of unwanted action. The safest tool is often not a general shell or database connection but a narrow function with validated parameters.
What Zero Trust Means for an AI Agent
Traditional zero trust is summarized as "never trust, always verify." Network location is not proof of authorization. Every request needs identity, policy evaluation, and minimal access. The principle fits agents because the model's decision process is probabilistic and exposed to untrusted context.
An agent should not receive trust simply because it runs inside the company's cloud account. A tool call should prove which agent made it, which user authorized the session, what resource is affected, and whether the action fits current policy. The system should assume credentials may be stolen, prompts may be injected, and generated code may be hostile.
Zero trust does not mean the model is useless or constantly blocked. It creates a safe operating envelope. The model can reason creatively inside that envelope, while the infrastructure enforces hard limits on identity, data, network access, transactions, and code execution.
Google's design uses three reinforcing layers: signed writes, sandboxed execution, and deterministic gateways. No single layer is sufficient.
Layer 1: Cryptographically Sign Every Write
Many agent systems use a shared database credential. Several worker agents connect through the same pool, and the database sees all writes as coming from the same service. If a row is changed, the team may not be able to prove which agent or session produced the action.
Google proposes giving each agent a cryptographic identity and requiring it to sign every state-changing payload. A refund record might include the agent ID, user or session ID, order ID, amount, timestamp, and policy version. The database ingress layer verifies the signature before committing the write.
In production, Google suggests a hardware-backed key in Cloud KMS or Cloud HSM. The private key is not stored as a plain environment variable inside the agent container. The agent receives permission to request a signature, while the key material remains protected by the key-management service.
The public key can be used to verify the signature. If an attacker changes the refund amount in the database after the write, the signature no longer matches the payload. An independent audit process can scan the ledger and identify tampering.
What Signed Writes Provide
Signed writes provide attribution, integrity, and non-repudiation within the system's trust model. A record can be connected to a specific agent identity. Unauthorized modification becomes detectable. Key permissions can be revoked without rewriting the entire application.
They also improve incident investigation. Instead of asking which shared service account changed a row, investigators can trace the action to an agent, key version, session, and policy context.
What Signed Writes Do Not Provide
A valid signature does not prove that the action was wise or authorized by the business. A compromised agent with legitimate signing permission can still sign a harmful refund. The database must validate both the signature and the policy.
Keys also require lifecycle management. Each agent should not automatically receive permission to sign every type of action. Roles need narrow scope. Keys need rotation, revocation, monitoring, and rate limits. A signature system without secure identity governance creates impressive evidence of a mistake rather than preventing it.
Layer 2: Sandbox Generated Code
Agents often generate code for calculations, data transformations, tests, or automation. Running that code directly with `exec()` or inside a loosely configured container is dangerous. The code may contain an accidental infinite loop, destructive file operation, secret-reading command, or network request to an attacker.
Google's pattern uses gVisor, an application kernel that runs in user space and creates a stronger isolation layer between the workload and the host operating system. It intercepts system calls instead of exposing the sandboxed application directly to the host kernel in the same way as a standard container.
The reference configuration also removes network access, drops Linux capabilities, limits memory and CPU, mounts the generated file read-only, and applies a short timeout. Each control matters. A sandbox with unrestricted internet can still leak data. A sandbox with no resource limit can be used for denial of service. A read-write host mount can damage files.
Why Ordinary Containers May Not Be Enough
Containers are excellent deployment tools, but many container configurations share the host Linux kernel. A kernel vulnerability or dangerous capability can create a path out of the container. Misconfigured mounts, privileged mode, or exposed sockets can turn a small mistake into host compromise.
gVisor adds another boundary by implementing a Linux-like interface in a user-space application kernel. This reduces direct exposure to host kernel attack surfaces. It is not a guarantee against every vulnerability, but it creates defense in depth.
Sandboxing Has a Cost
Stronger isolation can reduce compatibility and performance. Some system calls or advanced workloads may not behave exactly as they do on a normal Linux host. Teams need to test their workload and choose the right boundary for the risk.
The answer is not to remove the sandbox whenever a script fails. Instead, decide whether the agent truly needs general code execution. Many tasks can be implemented as narrow, audited functions that are safer and faster than arbitrary code.
Layer 3: Deterministic Semantic Gateways
A semantic gateway sits between the model and sensitive systems. It inspects input, output, and tool calls using deterministic policy. Despite the word "semantic," the final enforcement should not depend only on another generative model making a subjective judgment.
For a refund agent, the gateway can verify that the amount does not exceed the order total, the order belongs to the current customer, the refund has not already been issued, and the user has the necessary role. It can block credit-card numbers or API key patterns from leaving the system. It can reject direct SQL generated by the model and require a typed refund function instead.
Google's example includes checks for prompt-injection phrases, secret patterns, and transaction bounds. In production, exact phrase matching is only one signal; attackers can reword instructions. Stronger controls come from structured schemas, allowlisted actions, verified values from trusted systems, and explicit authorization.
Validate Before and After the Model
Input validation removes or labels untrusted content and rejects requests outside the application's scope. Output validation ensures the model has produced the required structure. Tool-call validation checks the intended action against current policy. Database validation confirms the final state change.
This creates multiple opportunities to stop an attack. If the model misunderstands the prompt, the tool boundary still blocks the action. If a tool-call validator has a bug, the database rule can enforce the transaction limit.
Test Gateways Like Software
Google recommends regression tests in CI/CD. Security policies should be expressed as contracts with test cases: valid refund allowed, excessive refund blocked, unknown agent rejected, secret exfiltration blocked, duplicate refund rejected, and model upgrade behavior reviewed.
Every prompt change and model migration can be tested against the same attack set. The test suite becomes institutional memory. It prevents a later developer from removing a rule because a new prompt appears more reliable.
Why the Three Layers Need Each Other
Cryptographic signatures identify the actor and reveal tampering, but they do not contain unsafe code. Sandboxing contains code, but it does not decide whether a refund is valid. A gateway enforces business rules, but it may not prove which agent created a record or prevent arbitrary code from attacking the host.
Together, the layers form a stronger system. The gateway approves a narrowly defined action. The agent signs the exact payload. The ingress service verifies signature and policy. Any generated computation runs in isolation. Logs and audits can reconstruct what happened.
This is defense in depth. It assumes one control can fail and ensures the failure does not automatically become a disaster.
Least Privilege: The Foundation Beneath the Layers
Each agent should have the smallest set of tools and permissions needed for the current role. A refund agent does not need access to payroll. A research agent does not need deployment credentials. A coding agent reviewing a pull request does not need permission to delete the production database.
Permissions should be scoped by resource, action, value, and time. A tool can permit refunds only for verified orders, below a limit, within a short session, and after customer authentication. Short-lived credentials reduce the damage of leakage.
Separate read from write. Many tasks can be completed with read-only access and a proposed action for human review. The most sensitive tool should not be loaded into the agent session unless it is needed.
If multiple agents collaborate, identity becomes even more important. Our guide to A2A vs MCP explains how agents and tools can communicate. A protocol transports a request; authorization determines whether the request should succeed.
Human Approval for High-Impact Actions
Autonomy should match risk. A low-value, reversible action may run automatically. A high-value payment, deletion, account lock, legal submission, public statement, or production deployment should often require human approval.
Approval must be meaningful. Show the reviewer the source data, proposed action, affected resource, reason, and any policy warning. A button labeled "approve" beside a vague summary encourages rubber-stamping.
Set thresholds. A support agent might automatically refund a small shipping fee but require approval for a full order. A coding agent might create a branch and run tests but require a maintainer to merge. A marketing agent might draft posts but not publish to every account.
Human approval is not a replacement for technical controls. Reviewers can make mistakes or approve malicious content. The deterministic transaction and permission rules should still apply.
Secure Tool Design
Avoid giving an agent a raw database connection when a narrow service can perform the task. Define a typed tool such as `request_refund(order_id, amount, reason)`. The service verifies authentication, fetches the trusted order amount, checks policy, records the request, and returns a structured result.
Do not let the model provide authoritative values that already exist in a trusted system. If the order total is in the database, fetch it. Do not accept the amount from a customer message. If the user's role is in the identity provider, verify it there rather than asking the model to infer it.
Separate data from instructions. Retrieved webpages, emails, and documents should be clearly marked as untrusted content. Tools should never interpret a string from a document as permission to call another tool.
Protect secrets. An agent may need to use a credential without reading it. A secret manager can inject authorization at the tool layer. The model's context should not contain raw API keys.
Session, Memory, and Cross-User Risk
Agents maintain session state so they can continue a task. That state can create data leakage if user identities are mixed, caches are reused incorrectly, or memory stores private facts without proper scope.
Bind every session to an authenticated principal. Tag state by tenant, user, project, and retention policy. Do not rely on the model to remember which customer owns a record. Enforce that relationship in the storage query.
Long-term memory should be opt-in and reviewable. Store concise facts rather than full sensitive conversations when possible. Give users and administrators deletion controls.
When an agent delegates to another agent, pass only the necessary context. A specialist that checks an invoice total does not need the customer's entire support history. Context minimization improves both privacy and reasoning quality.
Logging, Monitoring, and Audit
Record tool calls, authorization results, policy decisions, key versions, model versions, prompt versions, and state changes. Avoid placing secrets or unnecessary personal data in logs. An audit record should explain an action without becoming a second unprotected database.
Monitor for unusual behavior: repeated blocked refunds, rapid tool calls, attempts to access unrelated resources, sudden increases in errors, or a new model version producing different action patterns. Rate limits can stop a small failure from scaling.
Create an emergency stop. Operators should be able to disable a tool, revoke a key, pause an agent, or isolate a tenant without redeploying the whole platform. Test the stop procedure before an incident.
Enterprise stories such as how NVIDIA uses ChatGPT Work demonstrate the value of connected AI workflows. The security requirement grows with the value: once the system influences real decisions, auditability becomes part of the product.
A Step-by-Step Implementation Roadmap
Phase 1: Map Actions and Risk
List every tool and every state-changing action. Mark whether it affects money, private data, credentials, external communication, availability, or compliance. Identify which actions are reversible.
Phase 2: Narrow the Tools
Replace raw shell, database, and general HTTP access with typed functions. Validate parameters and fetch authoritative data inside the tool service.
Phase 3: Add Identity and Signatures
Give agents separate service identities. Sign sensitive payloads and verify them at ingress. Store agent, session, user, policy version, and key version in the audit record.
Phase 4: Isolate Code
Remove arbitrary code execution unless necessary. Where it is required, use strong sandboxing, no default network access, read-only mounts, resource limits, and timeouts.
Phase 5: Build Deterministic Policy
Implement transaction limits, role checks, tenant isolation, data-loss prevention, schema validation, and approval thresholds outside the model.
Phase 6: Test Attacks and Failures
Create a regression suite containing direct and indirect prompt injection, secret requests, unauthorized records, duplicate actions, malformed inputs, long content, tool errors, and model refusal. Run it after prompt and model changes.
Phase 7: Operate and Improve
Monitor blocks and false positives. Rotate keys. Review permissions. Update the threat model when tools are added. Practice incident response and confirm that backups can restore state.
Does This Apply Only to Google ADK?
No. Google's reference implementation uses Gemini, ADK, Cloud KMS, Cloud HSM, gVisor, and Google Cloud services, but the architectural ideas are portable.
AWS or Azure key-management systems can sign actions. A dedicated HSM can protect keys on premises. Other sandboxes, microVMs, or restricted execution services can contain code. An API gateway or policy engine can enforce schemas and business rules. Any agent framework can call narrow tools.
The important requirement is independence from the model. If the model can rewrite or bypass the policy through the same prompt channel, the policy is not a hard boundary.
Developers choosing a coding workflow can compare the control models in our Claude Code vs OpenCode guide. Local, cloud, and integrated agents expose different attack surfaces, but all need constrained credentials and review.
Tradeoffs and Costs
Zero-trust architecture adds engineering work. Key management, signatures, sandboxes, gateways, tests, and audit systems cost money. They can add latency and cause false positives. Strong isolation may break a dependency. Human approval can reduce automation speed.
The correct comparison is not secure automation versus free automation. It is the cost of controls versus the expected cost of unauthorized payments, data breaches, outages, legal exposure, and lost trust.
Teams can apply controls proportionally. A personal note organizer needs less infrastructure than a finance agent. Start with the highest-impact actions and protect them first. Do not connect an agent to a sensitive tool merely because the integration is easy.
Model capability also affects the design. Our GPT-5.6 vs Gemini 3.7 Flash vs Claude Opus 5 comparison helps with model selection, but a stronger model does not remove the need for hard boundaries. Greater capability can increase both usefulness and the consequences of excessive permission.
Common Security Mistakes
The first mistake is placing every rule in the system prompt. Move deterministic policy into code and infrastructure.
The second is using one powerful service account for all agents. Separate identities and grant narrow roles.
The third is executing generated code on the host or in a privileged container. Treat it as untrusted.
The fourth is allowing arbitrary URLs or network egress. Use allowlists, proxies, and no-network defaults.
The fifth is logging full prompts, responses, and secrets. Redact and minimize operational data.
The sixth is testing only normal user requests. Attackers, malformed inputs, and tool failures define the real security boundary.
The seventh is assuming a signature means the action was authorized. Verify both cryptographic identity and current business policy.
What Small Businesses Should Do
A small business may not need an HSM on day one, but it should still use the pattern. Keep the agent read-only until the workflow is proven. Require approval for refunds and payments. Use a dedicated account with limited permissions. Never paste passwords into prompts.
Choose platforms that provide audit logs, spending limits, secret storage, backups, and role controls. Turn on multi-factor authentication. Review every integration and remove unused access.
Begin with tasks that produce drafts rather than final actions: prepare a reply, summarize an order, suggest a refund, or create a report. A human can approve the result. Increase autonomy only after failure modes are understood.
Learning to supervise AI systems will remain valuable even as tools improve. Our guide to AI skills that will still matter in 2030 highlights judgment, verification, and workflow design - exactly the abilities needed to operate agents safely.
Final Thoughts
Google's zero-trust agent pattern begins with a realistic assumption: the model can be tricked. That does not make agents unusable. It tells developers where trust should live.
Sensitive writes should carry cryptographic identity and be verified before commitment. Generated code should run inside strong isolation with minimal resources and no unnecessary network. Inputs, outputs, and tool calls should pass through deterministic policies that the model cannot talk its way around.
Those three layers become stronger when combined with least privilege, human approval, scoped memory, secure tools, careful logs, regression testing, and a tested emergency stop. The goal is not to eliminate every mistake. It is to stop one mistaken or malicious instruction from becoming an unlimited production action.
As AI moves from answering to doing, this architecture will become ordinary software engineering. The companies that benefit most from agents will not be the ones that grant the most autonomy first. They will be the ones that create the clearest boundaries, the best evidence, and the safest path from a model's proposal to a real-world action.
Frequently Asked Questions
What is a zero-trust AI agent?
It is an agent deployed under the assumption that the model, prompt, tool, credential, or environment can be compromised. Every sensitive action is independently authenticated, authorized, constrained, and audited.
Why is a system prompt not enough?
System prompts are interpreted by a probabilistic model and can be affected by prompt injection, context, tuning, and model changes. Deterministic rules such as transaction limits should be enforced outside the model.
What are Google's three main security layers?
The reference pattern uses cryptographic signatures for state-changing writes, gVisor sandboxing for generated code, and deterministic semantic gateways for input, output, and tool-call policy.
Does signing an action make it safe?
No. It proves which key signed the payload and helps reveal tampering. The system must still verify that the action is authorized and within policy.
Is gVisor completely secure?
No technology eliminates all risk. gVisor adds a strong user-space isolation layer and reduces direct host-kernel exposure, but teams still need patched systems, restricted networking, resource limits, and monitoring.
Can this architecture work outside Google Cloud?
Yes. Equivalent key management, sandboxing, policy engines, and typed tools can be implemented on other clouds or on premises.
Should every agent action require human approval?
Not necessarily. Approval should match impact and reversibility. Low-risk actions can be automated, while high-value payments, deletions, public communications, and production changes often deserve review.
Official sources & references
Sources checked on 31 August 2026. Product features, availability and pricing can change; verify the linked primary source before acting.
- Official security announcement: Google: Beyond Zero for enterprise security
- Official framework: Google Secure AI Framework (SAIF)
0 Comments