Updated October 2, 2026.
Building an AI agent demo is easy. Building a multi-tenant AI agent platform that can safely act for thousands of customers is a very different engineering problem.
The model is only one component. A production platform also needs tenant isolation, delegated identity, tool authorization, human approvals, durable state, observability, evaluations, usage metering, billing controls, and an immutable audit trail. If any of these controls live only inside a prompt, they are not real controls.
This guide presents a provider-neutral production architecture for SaaS teams building agents that read data, call APIs, update records, run code, use MCP servers, and perform other actions with real side effects.
The Core Principle: The Model Plans, the Platform Decides
An LLM can propose an action, but it should never be the final authority on whether that action is allowed. Treat the model as an untrusted planner operating inside a trusted control plane.
The agent may request an action. Your deterministic policy layer must authorize it.
This separation prevents a prompt injection, hallucination, compromised tool response, or incorrect model decision from silently becoming a production incident.
A Reference Architecture
User / API Client | v Identity Gateway ----> Tenant Resolver ----> Entitlement Service | | v v Agent Control Plane ------------------------> Policy Engine | | | allow / deny / require approval v | Agent Runtime <--------- Approval Service <--------+ | +----> Model Gateway +----> Memory / Retrieval +----> Tool Gateway ----> SaaS APIs / MCP / Databases +----> Sandbox Executor | +----> Trace Pipeline +----> Usage Meter +----> Audit Ledger +----> Evaluation Pipeline
The architecture is easier to reason about when split into four planes:
-
Control plane: identity, policy, permissions, approvals, budgets, agent versions, and configuration.
-
Execution plane: model calls, tools, sandboxes, background jobs, retries, and resumable runs.
-
Data plane: tenant-scoped application data, memory, files, vectors, and generated artifacts.
-
Observability plane: traces, logs, metrics, usage events, audit events, and evaluation results.
Keep these boundaries explicit. For example, a sandbox may execute generated code, but it should not own billing policy, approval state, or long-lived credentials.
1. Design Tenant Isolation Before Agent Logic
Every request, run, tool call, memory query, usage event, and audit event must belong to a tenant. Tenant identity should be resolved once at the trusted ingress boundary and propagated through typed context—not copied from model-generated arguments.
A practical tenant context may contain:
{ "tenant_id": "ten_01J...", "actor_id": "usr_01J...", "actor_type": "user", "agent_id": "agt_01J...", "agent_version": 17, "session_id": "ses_01J...", "run_id": "run_01J...", "delegation_id": "dlg_01J...", "region": "us-east", "data_classification": "confidential" }
Use the tenant ID in every storage key and database access path. In PostgreSQL, include tenant_id in primary or unique constraints where appropriate, and consider Row-Level Security as defense in depth. Application-level scoping is still necessary: RLS should not become an excuse to pass unsanitized SQL or overly broad database credentials to an agent.
Recommended isolation layers
Layer
Minimum control
Higher-assurance option
Database
Tenant-scoped queries
RLS, separate schemas, or separate databases
Object storage
Tenant-prefixed keys
Per-tenant buckets or scoped access points
Vector search
Mandatory tenant filter
Separate indexes for sensitive workloads
Cache
Tenant in every key
Separate clusters for regulated tiers
Sandbox
Per-run workspace
Per-tenant VM or microVM isolation
Secrets
Server-side secret references
Per-tenant vault namespaces and short-lived credentials
2. Model Identity, Delegation, and Agent Versions
A production agent is not simply “the current user.” It is a delegated actor operating for a user, service account, or organization under a specific policy and for a limited duration.
Represent these identities separately:
-
Human identity: the person who requested or approved the work.
-
Service identity: the backend workload running the agent.
-
Agent identity: the configured agent and immutable version executing the run.
-
Tool identity: the credential or connection used to access an external system.
-
Delegation: the explicit link describing who authorized the agent to do what, for which tenant, and until when.
Never let an agent silently inherit every permission of the human who opened the chat. Mint a narrow delegation with a short expiry, approved scopes, resource boundaries, and a revocation path.
Version everything that changes behavior
Store an immutable version for instructions, model configuration, tool catalog, policy bundle, output schema, memory strategy, and routing rules. Every run should reference the exact versions used. Without versioning, you cannot reliably reproduce failures or compare evaluations.
3. Separate Entitlements from Authorization
Billing plans and permissions answer different questions:
-
Entitlement: Has this tenant purchased access to the feature or capacity?
-
Authorization: Is this actor allowed to perform this action on this resource right now?
A Pro plan may include the “CRM Agent” feature, but only a sales manager may be allowed to export contacts. A tenant may have unused credits, but a destructive tool may still require approval.
Evaluate both layers before execution:
`if !entitlements.Enabled(tenant, "crm_agent") { return FEATURE_NOT_AVAILABLE }
decision = policy.Evaluate({ tenant, actor, agent_version, tool, operation, resource, arguments, risk, budget, environment })
switch decision.effect { case "allow": execute() case "require_approval": pause_and_request_approval() default: deny() }`
The policy engine should return a structured decision with a reason, matched rule, policy version, required conditions, and expiry. Persist that decision with the run.
4. Build a Tool Gateway, Not a Bag of Functions
Tools are where an agent crosses from language into side effects. Every tool invocation should pass through a trusted gateway that owns validation and enforcement.
The gateway should:
-
Resolve the canonical tool by an internal ID—not a model-supplied URL.
-
Validate arguments against a strict schema.
-
Inject tenant and actor context server-side.
-
Evaluate permissions and risk.
-
Check budgets, quotas, and rate limits.
-
Request approval when required.
-
Acquire short-lived credentials.
-
Execute with timeout, network, and data boundaries.
-
Normalize and classify the result.
-
Emit trace, usage, and audit events.
Assign a risk level to every tool operation
Risk
Examples
Default treatment
R0 — Read-only
Search public docs, read a non-sensitive record
Allow with logging and rate limits
R1 — Sensitive read
Read customer data, financial data, or private files
Scope checks, redaction, enhanced audit
R2 — Reversible write
Create a draft, add a label, schedule a tentative event
Policy check; approval for unusual scope
R3 — External side effect
Send email, publish content, deploy code, issue a refund
Preview and explicit approval by default
R4 — Destructive or privileged
Delete data, rotate keys, transfer funds, change permissions
Strong approval, step-up authentication, or prohibit
Tool metadata is useful for classification, but metadata is not enforcement. An untrusted tool server can mislabel an operation. Your gateway must enforce network restrictions, credentials, scopes, and approvals independently.
MCP in a multi-tenant platform
If you expose or consume tools through the Model Context Protocol, keep authorization at the transport and gateway boundaries. Validate the token issuer, audience or protected resource, expiry, scopes, and tenant binding. Do not forward a user token broadly across unrelated tool servers. Prefer token exchange or narrowly scoped, short-lived credentials for each destination.
Keep tool discovery separate from tool authorization. Seeing a tool in a catalog must not imply permission to invoke it.
5. Guardrails Must Exist at Multiple Boundaries
“We added a safety prompt” is not a production security design. Use different controls for different boundaries:
-
Input guardrails: detect disallowed requests, sensitive data, unsupported tasks, and obvious injection patterns.
-
Retrieval guardrails: enforce tenant filters, document ACLs, source allowlists, and content classification.
-
Tool guardrails: validate arguments, destinations, amounts, resource IDs, and side-effect risk.
-
Output guardrails: validate schemas, redact secrets, detect unsupported claims, and enforce channel policy.
-
Runtime guardrails: cap turns, tools, tokens, time, concurrency, and total cost.
-
Infrastructure guardrails: sandboxing, egress allowlists, filesystem boundaries, scoped credentials, and process limits.
Fail closed when a mandatory policy service is unavailable. A timeout in the authorization layer should not become permission to continue.
6. Treat Human Approval as a Durable State Machine
Approval is not a modal dialog attached to a chat message. It is a durable workflow state.
RUNNING | v AWAITING_APPROVAL ---- reject ----> REJECTED | approve v READY_TO_EXECUTE ---- failure ----> FAILED | execute v COMPLETED
The approval record should include:
-
The exact tool, operation, destination, and normalized arguments.
-
A human-readable preview of the side effect.
-
The tenant, actor, agent version, run, and policy decision.
-
Who may approve and whether step-up authentication is required.
-
An expiry time and a hash of the approved payload.
-
The final decision, approver identity, timestamp, and reason.
If arguments change after approval, invalidate the approval. Resume the same durable run from stored state instead of starting a fresh agent turn that may produce a different action.
7. Keep Memory Scoped, Typed, and Deletable
“Memory” usually combines several different stores:
-
Conversation state: recent turns and resumable run state.
-
Working memory: temporary plans, summaries, and intermediate artifacts.
-
Semantic memory: retrieved facts, documents, and embeddings.
-
Profile memory: stable user preferences explicitly allowed for reuse.
-
Operational state: task handles, checkpoints, approval state, and job progress.
Each type needs a retention policy, tenant boundary, provenance, deletion path, and access policy. Do not put secrets into prompts or vector stores merely because retrieval is convenient. Store secret references and resolve them only inside the trusted execution boundary.
Attach provenance to retrieved content so the runtime knows the source, tenant, ACL, timestamp, and trust level. Treat retrieved text and tool output as untrusted data, not instructions.
8. Trace Every Run Across Models, Tools, and Approvals
A chat transcript is not enough to debug an agent. You need an end-to-end trace that captures decisions and side effects.
Use stable identifiers:
-
tenant_id -
session_id -
run_id -
trace_id -
step_id -
tool_call_id -
approval_id -
usage_event_id
Useful spans include model inference, retrieval, policy evaluation, approval wait, tool execution, sandbox execution, background jobs, and final response generation.
Record duration, status, model, token usage, tool name, retry count, policy decision, and error class. Avoid recording raw secrets, credentials, or unrestricted customer payloads. Store hashes or redacted summaries where full inputs are unnecessary.
Operational metrics that matter
-
Run success and task completion rate.
-
Latency by model, tool, and workflow step.
-
Tool selection and argument-validation failures.
-
Approval rate, rejection rate, and approval wait time.
-
Retry, timeout, and circuit-breaker rate.
-
Input, output, cached, and reasoning-token usage where available.
-
Cost per successful task—not merely cost per model call.
-
Tenant-level error, quota, and spend anomalies.
9. Make Evaluations Part of Deployment
Traditional unit tests remain essential, but they do not measure whether an agent selected the right tool, followed the correct path, or completed the user’s goal.
Build an evaluation pyramid:
-
Deterministic tests: schemas, policy rules, tenant isolation, billing math, idempotency, and state transitions.
-
Tool-contract tests: valid and invalid arguments, permissions, timeouts, retries, and normalized errors.
-
Golden-task evals: representative user goals with expected outcomes and forbidden actions.
-
Trace evals: grade tool choice, handoffs, policy compliance, efficiency, and recovery behavior.
-
Adversarial evals: prompt injection, cross-tenant requests, secret extraction, approval bypass, and malicious tool output.
-
Online monitoring: sampled production traces, user corrections, support incidents, and business outcomes.
Run the regression suite whenever you change a model, prompt, tool schema, policy, memory strategy, routing rule, or agent version. A model upgrade is a behavioral software change and should pass the same release discipline as application code.
Example release gates
Metric
Example gate
Cross-tenant data access
Zero tolerated failures
Forbidden tool execution
Zero tolerated failures
Task completion
No statistically meaningful regression
Approval bypass
Zero tolerated failures
Cost per successful task
Within the product budget
P95 latency
Within the workflow SLO
10. Build Usage Metering as an Append-Only Ledger
Do not calculate customer bills by scanning application logs at the end of the month. Emit normalized, idempotent usage events as work occurs.
{ "event_id": "use_01J...", "tenant_id": "ten_01J...", "run_id": "run_01J...", "meter": "agent_tokens", "quantity": 18432, "unit": "token", "provider": "model-provider", "model": "model-name", "occurred_at": "2026-10-02T12:00:00Z", "idempotency_key": "run_...:model_step_7", "dimensions": { "agent_id": "agt_01J...", "environment": "production" } }
Useful meters may include model tokens, model requests, tool executions, sandbox seconds, vector operations, storage, background task time, and premium connector usage.
Reserve, meter, settle
-
Reserve: estimate the maximum cost and verify quota before starting expensive work.
-
Meter: emit usage events during execution.
-
Settle: release unused reservation and post the final amount.
This pattern prevents a single runaway run from consuming an entire tenant budget. Add per-run, daily, monthly, and organization-level limits. Support hard stops, soft alerts, and administrative overrides as distinct policies.
Keep provider cost, customer usage, and customer price as separate concepts. That allows markup, credits, bundles, free allowances, and pricing changes without rewriting historical events.
11. Create an Immutable Audit Trail
Traces help engineers understand how a run behaved. Audit logs help security, compliance, administrators, and customers understand who did what and why.
Capture audit events for:
-
Agent creation, configuration, version changes, and deployment.
-
Tool installation, connection, scope, and credential changes.
-
Policy and entitlement changes.
-
Sensitive reads and all material writes.
-
Approval requests, approvals, rejections, and expirations.
-
Exports, deletions, permission changes, and billing overrides.
-
Administrative impersonation or support access.
A strong event contains the actor, tenant, action, resource, time, source, agent version, policy decision, approval reference, outcome, and correlation IDs. Use append-only storage, strict write permissions, retention policies, integrity controls, and export support for enterprise SIEM systems.
Do not confuse “immutable” with “store every raw payload forever.” Minimize sensitive data, redact secrets, hash large payloads when appropriate, and support legally required deletion workflows without destroying the integrity of unrelated audit records.
12. Reliability: Assume Every Dependency Will Fail
Agent workflows combine probabilistic models with APIs, databases, queues, and external SaaS systems. Partial failure is normal.
Design for:
-
Idempotency: retries must not send the same email, refund, or deployment twice.
-
Timeouts: set deadlines for models, tools, approvals, and jobs.
-
Retry policy: retry only transient failures with capped exponential backoff and jitter.
-
Circuit breakers: stop cascading failures from unhealthy providers.
-
Checkpoints: persist durable state after meaningful steps.
-
Compensation: define how to reverse or reconcile partially completed workflows.
-
Dead-letter handling: preserve failed jobs for safe inspection and replay.
-
Concurrency control: prevent two runs from modifying the same resource incompatibly.
Use a stable idempotency key derived from the tenant, run, step, tool, and intended operation. A retry should reuse the same key and payload. If the payload changes, create a new operation.
13. A Practical PostgreSQL Data Model
A production schema will vary, but these tables form a useful starting point:
tenants users memberships service_accounts agents agent_versions agent_delegations tool_definitions tool_connections tool_permissions entitlements policy_bundles sessions agent_runs run_steps approval_requests usage_events budget_reservations audit_events eval_datasets eval_runs eval_results
Make event and ledger tables append-only at the application layer. Use unique idempotency keys for tool effects and usage events. Partition high-volume trace, audit, and usage tables by time, while keeping tenant-aware indexes for operational queries.
14. End-to-End Request Lifecycle
-
The gateway authenticates the caller and resolves the tenant.
-
The platform loads an immutable agent version.
-
The entitlement service verifies feature access and available capacity.
-
The budget service reserves an upper cost bound.
-
The runtime starts a run and trace.
-
The model proposes a plan or tool call.
-
The tool gateway validates the schema and injects trusted context.
-
The policy engine returns allow, deny, or require approval.
-
If approval is required, the run pauses with durable resumable state.
-
After approval, the gateway obtains narrow, short-lived credentials.
-
The operation executes inside network, time, and resource boundaries.
-
The platform records the result, trace spans, usage, and audit event.
-
The output guardrail validates and redacts the response.
-
The budget reservation is settled against actual usage.
-
Selected traces enter evaluation and quality-monitoring pipelines.
15. Production Checklist
Identity and tenancy
-
Every request and record has a trusted tenant context.
-
Agents use explicit, expiring delegations.
-
Credentials are short-lived, scoped, revocable, and never placed in prompts.
Tools and safety
-
Every tool has a schema, owner, risk level, timeout, and idempotency policy.
-
All side effects pass through deterministic authorization.
-
Sensitive actions have previews and durable approvals.
-
Sandboxes have restricted filesystems, networks, credentials, and compute.
Observability and quality
-
Runs are traceable across models, tools, approvals, and jobs.
-
Logs and traces redact secrets and minimize personal data.
-
Model, prompt, policy, and tool changes trigger regression evals.
Billing and compliance
-
Usage events are append-only and idempotent.
-
Budgets exist at run, tenant, and billing-period levels.
-
Audit events capture administrative and agent actions.
-
Retention, deletion, export, and regional storage policies are documented.
Final Takeaway
The winning multi-tenant agent platforms will not be the ones with the longest prompts. They will be the ones with the strongest control planes.
Build the platform so that identity is explicit, tenant boundaries are enforced outside the model, tools receive least privilege, sensitive actions pause for approval, every run is traceable, changes are evaluated, usage is metered, and important actions are auditable.
Once those foundations exist, models and tools can evolve without forcing you to rebuild trust, security, and economics from scratch.
Further Reading
-
OpenAI: Guardrails and human review
-
OpenAI: Evaluate agent workflows
-
OpenAI: Agent tracing
-
Model Context Protocol specification
-
OpenTelemetry semantic conventions for generative AI


Discussion (0)