All posts Security··Dane Holcombe, Founder, Orqana AI

Giving an AI Agent Write Access, Safely

Read-only agents are safe but limited. The moment an agent can send email, update a record, or create a calendar invite, it becomes genuinely useful — and genuinely risky. Here's the access model we landed on, and why each piece of it matters.

The write access problem

Most AI agent demos show read-only operations. The agent summarizes an email thread. It looks up a record. It answers a question from a knowledge base. Safe, controllable, easy to demo.

The agents teams actually want to deploy do more than read. They draft and send replies. They update CRM records. They create tickets, schedule meetings, and file expense reports. That's where the value is — not in surfacing information, but in taking action.

Write access is what separates an AI research assistant from an AI operations assistant. But it's also what makes mistakes consequential. A read-only agent that gets something wrong produces a bad answer. A write-access agent that gets something wrong sends an email to the wrong person, overwrites a record, or creates a meeting that shouldn't exist.

The question isn't whether to give agents write access. The question is how to do it in a way where the benefits are real and the downside is bounded.

AI agent guardrails: what they actually are

"Guardrails" is a term that gets used loosely. In practice, AI agent guardrails are the combination of instructions, access controls, and logging that constrain what an agent can do and make its behavior auditable.

Guardrails operate at three levels:

Instruction-level guardrails. What the agent is explicitly told to do and not do in its brief. "Never send an external email without summarizing what you're sending and to whom in the activity log." "If a request involves account cancellation, escalate to a human with a summary rather than processing it." These are the agent's rules of engagement.

Integration-level guardrails. What the agent is technically capable of doing, based on which integrations are attached and with what scopes. An agent that doesn't have a CRM integration attached can't update CRM records, regardless of what someone asks it to do. Access is gated at the platform level, not just the instruction level.

Observability guardrails. The run log that records every action the agent took. Not as a punitive measure, but as the mechanism that lets you catch unexpected behavior before it becomes a pattern, and attribute actions to specific runs when you need to.

All three matter. Instruction-level guardrails can be overridden by adversarial inputs or edge cases the brief didn't anticipate. Integration-level guardrails provide a hard floor — the agent literally cannot do what it doesn't have access to do. And observability ensures that when something unexpected happens, you find out quickly.

The per-agent scoping model

The most important structural decision in safe AI agent deployment is per-agent integration scoping: each agent only has access to the integrations it explicitly needs for its defined task.

This sounds obvious, but it's not universal. Some platforms give workspace-level integration access — connect Gmail once, and every agent in the workspace can read and write to Gmail. That's convenient for setup, but it means a support triage agent, a sales research agent, and a financial reporting agent all have access to your email. If any one of them behaves unexpectedly, the blast radius is large.

Per-agent scoping contains the blast radius. Here's what that looks like in practice:

Support triage agent: access to the shared support inbox (read and write) and the internal knowledge base (read only). No access to CRM, financial systems, or calendar.

Sales research agent: access to email history (read only) and CRM (read only). No write access to anything — it produces summaries for human review, it doesn't take action.

IT onboarding agent: access to the directory service (read and write for provisioning) and the ticketing system (read and write). No access to email or financial systems.

Each agent is a narrow, bounded actor. If the support triage agent misbehaves, the worst case is something wrong happening in the support inbox — not across the entire organization's connected systems.

Orqana AI implements per-agent scoping as a core feature. Integration access is attached per agent, not at the workspace level. Disconnecting an integration revokes it from the agent immediately, with no downstream cleanup required.

Explicit escalation: the most underused guardrail

The single most effective guardrail for write-access agents is explicit escalation logic: the agent knows what it's empowered to do, and for everything outside that set, it escalates with a summary rather than attempting to handle it.

Most briefs don't include this. They describe what the agent should do, but don't describe what it should do when it encounters something it shouldn't handle. That gap is where most write-access incidents come from — not adversarial inputs, but legitimate requests that fall outside the agent's intended scope, handled by an agent that doesn't know it's out of its depth.

Good escalation logic in a brief looks like:

"For any request involving refunds over $500, do not process. Summarize the request, the customer's account history, and your recommendation, and route to the billing team with the subject line '[ESCALATE] Refund review needed'."

"If you're uncertain about the correct answer to a technical question, say so explicitly and offer to connect the customer with a specialist. Do not guess."

"Never send an email that includes specific pricing commitments, contract terms, or statements about product roadmap. Flag these to the sales team for review."

These instructions don't limit what the agent can do. They define the decision boundary between what it handles autonomously and what it hands off. A well-defined boundary makes write access safe — not by making the agent less capable, but by making it more self-aware about where its authority ends.

Testing write access before going live

Unlike read-only agents, write-access agents need to be tested against their actual integrations before they're deployed to production. Here's what a minimal pre-deployment test looks like:

Happy path testing. Run five to ten representative requests through the agent and verify that the actions taken are correct — the right email was drafted, the right record was updated, the right ticket was created.

Boundary testing. Send requests that should trigger escalation and verify that the agent escalates correctly rather than attempting to handle them. This is the most common gap in pre-deployment testing.

Adversarial input testing. Send inputs designed to make the agent take actions it shouldn't — requests that try to override instructions, requests that probe what happens at the edges of the agent's defined scope. If the agent handles these correctly, the guardrails are working. If it doesn't, fix the brief before deploying.

Volume testing. If the agent will handle high volume, run a batch of historical requests through it (with write actions sandboxed) to spot patterns. A single test case might pass while the agent systematically mishandles a class of inputs.

The goal of pre-deployment testing isn't to prove the agent is perfect. It's to verify that the failure modes are bounded and that escalation works as designed.

The run log as accountability infrastructure

Once a write-access agent is live, the run log becomes your most important operational tool. Every action the agent takes should appear in the log with enough context to reconstruct what happened and why.

A useful run log entry for a write-access agent includes:

  • The input that triggered the run
  • The agent's reasoning (what it decided and why)
  • The specific actions taken (what was written, to where)
  • The time and any relevant identifiers (the ticket ID, the email thread, the CRM record)

This isn't just for debugging. It's accountability infrastructure. When a customer asks "did your agent send me an email?" you can answer specifically. When an audit requires you to demonstrate what actions were taken and under what authority, the run log provides it. When a new employee asks "has the agent ever done X?", you can check.

Orqana AI's activity log captures this for every agent run. It's searchable by agent, by date, and by action type. For teams in regulated industries, it's the basis of the audit trail that compliance requires.

The trust model for write-access AI agents

Building organizational trust in a write-access AI agent is as much a process question as a technical one. Here's the sequence that works:

Start read-only. Even if you intend to give the agent write access, start with read-only permissions and validate the agent's behavior — its answers, its escalation decisions, its understanding of context — before it starts taking action.

Add write access incrementally. Grant write access to one integration at a time, starting with the lowest-stakes one. Add CRM write access before email write access. Add draft-and-review mode before send access.

Run in shadow mode. Before the agent acts autonomously, have it propose actions that a human reviews and approves. This builds the team's confidence in the agent's judgment and surfaces edge cases before they become incidents.

Set a review cadence. Establish a regular check on the run log — weekly at first, then as confidence builds, less frequently. The check isn't to find problems (though it will catch some) — it's to maintain the habit of oversight that makes write-access AI sustainable.

Expand scope as trust is earned. As the agent handles its initial scope reliably, expand it. An agent that has handled two hundred support escalations correctly has earned the trust required to handle two hundred and one.

This isn't a slow process. A well-built agent can move through this sequence in two to three weeks. The point is that the speed of trust-building is determined by evidence — by what the run log shows — not by optimism about what the agent should be able to do.

The brief as the primary guardrail

Everything else — integration scoping, escalation logic, the run log — is secondary to the brief itself. The brief is the agent's operating manual. If the brief is vague, no amount of technical guardrailing will make the agent safe. If the brief is precise, most failure modes don't arise.

A brief that makes write access safe typically includes:

  • A clear statement of what the agent is authorized to do autonomously
  • An equally clear statement of what requires human review or approval
  • Specific language about what information the agent must never include in external communications
  • A defined escalation path for requests outside the agent's scope
  • A tone and communication style guide for any written outputs

This isn't a long document. A well-written brief can be two to four paragraphs. The length isn't the point — the precision is. Vague instructions produce vague behavior. Precise instructions produce precise behavior.

Orqana AI's templates include brief structures that have been tested in production for common write-access use cases. They're not a substitute for your own specification, but they're a faster starting point than a blank page. The security overview covers the technical access model in detail.

Describe your first agent today

Free to start. No credit card, no setup calls, no engineering ticket.

Get started free