Google Research Proposes Contextual Integrity Policy Engines and Multi-Agent Benchmarks for Agentic Privacy
A Google Research workshop report proposes LLM-driven policy engines based on contextual integrity, plus standardized multi-agent benchmarks, to make AI agents more privacy-aware, auditable and safer to deploy. What it means for data and AI engineers.

AI agents are becoming increasingly capable of doing things on our behalf.
They can read documents, access databases, call APIs, use external tools, communicate with other agents, and execute long-running workflows with considerably less human intervention than traditional software.
That capability creates a problem that traditional permission systems weren’t designed to solve.
An agent may have permission to access a piece of information, but that doesn’t necessarily mean it is appropriate for the agent to use or transmit that information in a particular context.
That distinction is at the center of a recent Google Research workshop report, Open and Emergent Problems in Agentic Privacy and Security: A Contextual Angle, introduced on the Google Research blog on October 5, 2026 by Eugene Bagdasarian and Marco Gruteser.
The report came out of the Google Contextual Agent Privacy and Security (CAPS) Workshop, held in New York in late 2025 with more than 50 academic and industry leaders. It explores how contextual integrity could provide a foundation for privacy and security policies in increasingly autonomous agentic systems. (Google Research)
It also calls for standardized environments for evaluating the safety of multi-agent systems.
The important part isn’t that Google has released a production policy engine.
It hasn’t.
The interesting part is the architectural direction being proposed:
Instead of asking only whether an agent is technically allowed to access information, ask whether that information flow is appropriate in the current context.
For engineers building systems that allow agents to interact with personal, enterprise, or sensitive data, that is a significant shift in thinking.

The proposal on one page: a contextual policy engine between agents and the data and tools they use, plus open “Agent Gym” environments to test how multiple agents behave together.
Why traditional permissions aren’t enough for AI agents
Traditional access control generally answers a relatively straightforward question:
Can this identity access this resource?
For example:
User → Database
│
└── SELECT permission? → YES
That works well when the software behavior is relatively predictable.
An autonomous agent is different.
Imagine an agent has legitimate access to a customer’s account information.
The agent could potentially:
- read the customer’s address
- send that address to another service
- include it in a generated report
- share it with another agent
- use it to make a decision
- store it in a different system
The fact that the agent can access the information doesn’t answer whether every one of those actions is appropriate.
The missing dimension is context.
Google Research’s report argues that traditional manual permissions and expert-written policies struggle to scale when agents can dynamically discover tools, interact with other agents, and execute long-running tasks.
That leads to a different question:
Is this information flow appropriate here, for this purpose, between these parties, under these conditions?
Contextual integrity gives us a different model
The idea comes from contextual integrity, a privacy theory developed by Helen Nissenbaum around the principle that privacy depends on whether information flows appropriately within a particular social context.
Instead of treating privacy as simply:
allowed / denied
the system considers the circumstances surrounding an information transfer.
The Google Research report applies this idea to agentic systems and discusses three important components:
Actors
+
Information type
+
Transmission principle
Actors
Who is involved?
For example:
User
Agent
Organization
Third-party service
Another agent
Information type
What information is being transferred?
For example:
Personal data
Financial data
Enterprise data
Documents
Operational data
Credentials
Transmission principle
Under what conditions is the information being transferred?
For example:
User consent
Specific task
Business purpose
Legal requirement
Security requirement
Together, these provide a much richer policy description than a simple access-control rule.
From access control to contextual policy
Consider an employee asking an AI agent:
“Find the latest status of my company’s supplier contracts.”
A traditional permission system might determine:
Employee
↓
Contract database
↓
Access allowed
A contextual policy engine could reason about something more specific:
Actor:
Employee + Agent
Information:
Supplier contracts
Purpose:
Contract-status investigation
Destination:
Internal workspace
Transmission:
Allowed for this task
Now change one variable.
The agent wants to send the same contract information to an external summarization service.
The underlying database permission hasn’t changed.
But the context has.
A contextual policy could therefore produce:
Internal retrieval → ALLOW
External transmission → BLOCK
That is the conceptual leap.
The policy isn’t merely attached to the database.
It evaluates the information flow.
What Google Research is proposing: a contextual policy engine
The report explores using LLMs to make these contextual policies machine-readable and adaptable to autonomous agents. It describes the contextual policy engine as part of a supervisor layer that monitors and enforces whether an agent’s actions are appropriate.
The proposed architecture can be thought of as a policy layer sitting between an agent and the systems it wants to access.
User request
│
▼
Autonomous agent
│
Proposed action
│
▼
┌─────────────────────────┐
│ Contextual policy engine│
│ │
│ Actors │
│ Information type │
│ Transmission principle │
└────────────┬────────────┘
│
┌──────┴──────┐
│ │
ALLOW BLOCK
│ │
▼ ▼
Tool / data Stop action
access
The key idea is that the policy decision happens in context, rather than being completely determined before the agent begins operating.
That matters because agents can encounter situations that weren’t explicitly anticipated when the original policy was written.
The dynamic policy problem
This is where agentic systems become especially difficult.
Traditional software often has a relatively predictable set of actions.
An autonomous agent might discover a new tool during execution.
It might call an API.
That API might return new information.
The agent might then decide that another tool is useful.
Another agent might become involved.
The information flow therefore evolves as the task evolves.
Google Research describes a dynamic policy-generation loop, which can run in real time, intended to evaluate these situations as they arise.
Conceptually:
Task starts
│
▼
Agent proposes action
│
▼
Context evaluated
│
▼
Policy generated / evaluated
│
├───────────────┐
│ │
ALLOW BLOCK
│ │
▼ ▼
Execute Reject
│
▼
New information / tool
│
▼
Re-evaluate context
This is fundamentally different from defining one giant permission file before the agent starts.
Why this matters for multi-agent systems
One agent is already difficult to govern.
Multiple autonomous agents make the problem significantly harder.
Imagine:
Supervisor
│
┌──────────┼──────────┐
│ │ │
Research Data Action
agent agent agent
│ │ │
└──────────┼──────────┘
│
External tools
Now information can move:
Agent A → Agent B
Agent B → Agent C
Agent C → External API
Agent A → Database
Agent C → User
The safety problem is no longer just:
“What can this agent access?”
It becomes:
“What happens when several agents collectively perform a sequence of actions that no individual agent was explicitly programmed to perform?”
That is the multi-agent safety problem Google Research is highlighting.
The benchmark problem
Current AI evaluations often focus on individual models.
For example:
Prompt
↓
Model
↓
Answer
↓
Score
That works reasonably well for evaluating a chatbot.
It is much less useful for evaluating an autonomous multi-agent system.
A multi-agent system behaves more like:
┌── Agent A ──┐
│ │
User → Supervisor ┼── Agent B ──┼→ Tools
│ │
└── Agent C ──┘
│
▼
New information
│
▼
More decisions
The resulting behavior can depend on:
- previous actions
- agent-to-agent communication
- tool outputs
- changing context
- accumulated state
- unexpected interactions
- failures and recovery behavior
A benchmark that evaluates only one model response cannot capture much of this.
Agent Gym: evaluating agents in environments
Google Research’s report calls for standardized, multi-agent benchmarks: dynamic “Agent Gym” environments where researchers can safely simulate complex, cascading interactions over extended periods, as open-source sandboxes.
The idea is important because researchers need environments that can reproduce complex interactions rather than simply comparing isolated model outputs.
A useful benchmark environment might look like:
Agent Gym
│
┌────────────┼────────────┐
│ │ │
Agent A Agent B Agent C
│ │ │
└────────────┼────────────┘
│
Shared environment
│
┌───────────┼───────────┐
│ │ │
Data Tools Users
The evaluation could then measure:
Privacy violations
Unauthorized actions
Information leakage
Agent cooperation
Tool misuse
Emergent behavior
Recovery from failures
Long-horizon safety
That is much closer to the environment in which production agents actually operate.
The engineering trade-off
There is an obvious cost to putting a policy evaluation layer in front of agent actions.
Every additional check introduces latency, potentially cost, and architectural complexity.
An agent that previously did:
Agent → Tool
now does:
Agent
↓
Policy engine
↓
Tool
For multi-step agents, that can happen dozens or hundreds of times.
So the engineering challenge becomes:
How much policy evaluation is necessary, and where should it happen?
A sensible architecture might eventually use different enforcement levels:
Low-risk action
→ lightweight policy check
Sensitive data
→ stronger contextual evaluation
External transmission
→ strict policy + approval
High-impact action
→ policy + human confirmation
The Google Research report does not provide a production implementation or performance benchmark for this architecture.
That is an important limitation.
The report is proposing a direction for research, not announcing a finished security product.
This is where data engineers should pay attention
Agentic privacy isn’t only an AI-model problem.
It’s also a data architecture problem.
If your agent has access to:
PostgreSQL
Data lake
Customer records
Vector database
Internal APIs
Cloud storage
Third-party SaaS
then the security boundary isn’t simply the model.
The data layer becomes part of the agent’s decision-making environment. That is one reason governed data platforms, like the lakehouse architecture we looked at with Databricks, matter so much for agents.
This means data engineers increasingly need to think about:
- data classification
- lineage
- identity
- purpose limitation
- access policies
- audit trails
- tool permissions
- data egress
- agent-to-agent communication
The traditional architecture:
Application → Database
is becoming:
Agent
│
├── Database
├── API
├── Vector search
├── SaaS
├── Another agent
└── External service
Every arrow is a potential information-flow boundary.
What AI engineers should do today
The proposed policy-engine architecture isn’t something you need to wait for.
There are several ideas you can apply to existing agent systems right now.
1. Treat tool access as a policy problem
Don’t give an agent unrestricted access simply because the underlying service supports it.
Define:
Agent identity
+
Tool
+
Data type
+
Purpose
+
Allowed destination
as explicit policy dimensions.
2. Log information flows, not just API calls
A normal application log might tell you:
Agent called customer_api
That’s useful, but insufficient.
A stronger audit record could capture:
Actor:
Customer-support agent
Data:
Customer account information
Purpose:
Resolve support ticket
Destination:
Internal CRM
Decision:
Allowed
Policy:
Customer-support-context-v3
That creates an audit trail that can actually explain why an action was allowed.
3. Build multi-agent tests
If you have three agents in production, don’t test them only individually.
Test interactions.
For example:
Research agent
↓
Data agent
↓
Action agent
↓
External API
Then deliberately introduce:
- conflicting instructions
- unexpected tool output
- sensitive information
- malicious prompts
- incorrect agent assumptions
- privilege escalation attempts
The goal is to discover what happens when the system behaves differently from the happy path.
4. Separate capability from authorization
An agent having access to a tool doesn’t mean it should use the tool in every context.
That’s an important distinction:
Capability
≠
Authorization
≠
Appropriateness
A mature agent architecture should eventually distinguish all three.
The biggest open questions
This proposal is promising, but there are still major unanswered questions.
Can an LLM reliably understand contextual norms?
This is probably the hardest problem.
Human privacy norms are often ambiguous.
For example:
“You can share my location with my delivery driver.”
Does that mean:
- this delivery only?
- this company permanently?
- any employee?
- third-party logistics providers?
- future deliveries?
Humans can resolve these distinctions using context.
Whether an LLM can reliably make those decisions under adversarial conditions is still an open research question.
What happens when policies conflict?
Imagine:
Company policy → ALLOW
User preference → BLOCK
Legal requirement → ALLOW
Third-party policy → BLOCK
Which policy wins?
A production policy engine needs deterministic conflict-resolution rules.
An LLM alone shouldn’t be expected to invent those rules at runtime.
What happens when the policy engine itself is attacked?
If the agent is autonomous and the policy engine is also LLM-driven, the system has another attack surface.
An attacker might attempt:
Prompt injection
↓
Agent
↓
Policy-engine manipulation
↓
Unauthorized action
That means the policy engine itself needs:
- isolation
- monitoring
- deterministic enforcement
- strong identity boundaries
- tamper-resistant logging
The proposal doesn’t eliminate the security problem.
It moves part of the security architecture into a new layer.
The bigger picture
What I find most interesting about this Google Research proposal is that it points toward a broader change in how we think about AI security.
Traditional security often asks:
Who are you, and what are you allowed to access?
Agentic security increasingly needs to ask:
What are you trying to do, what information are you using, why are you using it, who will receive it, and is that information flow appropriate in this context?
That’s a much harder problem.
But it may be necessary if autonomous agents are going to operate across enterprise systems rather than remaining isolated chat interfaces.
The architecture starts to look like:
USER
│
▼
AGENT SYSTEM
│
▼
CONTEXTUAL POLICY
EVALUATION
│
┌────────────┼────────────┐
│ │ │
Data Tools Agents
│ │ │
└────────────┼────────────┘
│
External world
And above all of this sits a broader evaluation layer:
Multi-agent safety benchmark
│
┌─────────┴─────────┐
│ │
Privacy testing Security testing
│ │
└─────────┬─────────┘
│
Agent behavior
The report frames this as defense at several levels at once: system, model, user, and ecosystem.
That is the direction worth watching.
The engineering takeaway
The future of agentic security probably won’t be solved by one permission system, one model, or one safety benchmark.
It will require several layers working together:
Identity
+
Data governance
+
Tool authorization
+
Contextual policy
+
Agent isolation
+
Runtime monitoring
+
Multi-agent evaluation
+
Human oversight
Google Research’s proposal is interesting because it connects two problems that are often discussed separately:
How should agents decide whether an information flow is appropriate?
and
How should we evaluate systems where multiple agents interact over long periods of time?
Contextual integrity offers one possible foundation for the first problem.
Agent Gym and standardized multi-agent environments point toward a possible approach for the second.
Neither is a solved production technology yet.
But as agents become more autonomous, these are exactly the kinds of problems that data and AI engineers will eventually have to solve.
The next generation of AI systems won’t just need to know what they can do. They will need to understand what they should do, and we will need reliable ways to test whether they actually do it.
Key takeaways
- Contextual integrity provides a framework for evaluating whether an information flow is appropriate within a specific context.
- Google Research proposes applying that idea to autonomous agents through LLM-driven, context-dependent policy evaluation in a supervisor layer.
- A contextual policy can reason about actors, information type, and transmission principles, rather than relying only on static permissions.
- Multi-agent systems create new safety challenges because interactions between agents can produce behavior that isn’t visible when evaluating agents individually.
- Google Research calls for “Agent Gym” and standardized multi-agent environments to improve reproducibility and evaluation of long-horizon agent behavior.
- These ideas remain research proposals rather than mature production technologies.
- For engineers today, the practical lessons are to treat tool access, data movement, information flows, and multi-agent interactions as explicit security boundaries.
- The long-term goal is not simply more restrictive agents. It is agents whose actions can be explained, evaluated, governed, and audited in context.
Sources
- Google Research Blog: Open and Emergent Problems in Agentic Privacy and Security: A Contextual Angle, Eugene Bagdasarian and Marco Gruteser (October 5, 2026)
- Google Research: the full technical report
- Helen Nissenbaum: Contextual Integrity Up and Down, the theory behind the proposal
Discussion
Join the discussion: create a free account or sign in to comment.
No comments yet.