Multi-Tenant SaaS Isolation Architecture: PostgreSQL RLS, Caches, Queues, Storage, and AI
Multi-tenancy is what makes SaaS economically attractive: many customers can share the same application and infrastructure.
It is also what makes one class of bug especially dangerous.
If Tenant A can ever read, modify, cache, process, retrieve, or export Tenant B's data, the problem is not a normal application bug. It is a broken tenant boundary.
That boundary does not stop at PostgreSQL. A SaaS can correctly filter SQL and still leak data through Redis, background workers, object storage, search, analytics, RAG, or AI tools.
The right mental model is simple:
> Tenant isolation is an end-to-end invariant across the whole request and data path.
The Core Architecture
Authenticated request
↓
Trusted tenant resolver
↓
Tenant context
↓
API / domain services
↓
PostgreSQL → tenant_id + RLS
Redis → tenant-scoped keys
Queues → tenant-scoped jobs
Storage → tenant-scoped objects
Search/RAG → mandatory tenant filters
Analytics → tenant attribution
AI tools → server-injected tenant
↓
Logs / traces / audit / usage
The tenant identifier should be resolved from trusted identity or machine credentials and propagated through server-side context.
It should not be freely chosen by the browser, background job payload, model, or tool argument.
Start by Defining the Tenant
A tenant can be a company, workspace, restaurant, store group, or enterprise account.
A tenant is usually not the same thing as a user.
One tenant can have many users, and a user can sometimes belong to multiple tenants.
A useful server-side context might contain:
tenant_id
actor_id
actor_type
membership_id
role
request_id
region
The important point is consistency. Every subsystem should receive the same trusted tenant identity.
Authentication Is Not Tenant Isolation
Authentication answers:
> Who are you?
Tenant isolation answers:
> Which customer's resources can this operation ever reach?
Authorization answers:
> Is this actor allowed to perform this action?
A user can be correctly authenticated and still request a resource belonging to another tenant.
That is why isolation must exist independently from login and normal role checks.
AWS's SaaS guidance treats tenant isolation as a foundational SaaS responsibility rather than something authentication provides automatically.
Choose Pool, Bridge, or Silo Deliberately
AWS describes three common SaaS partitioning models.
Pool
Tenants share infrastructure, database, and often schema.
orders
----------------
id
tenant_id
customer_id
total
Benefits:
- lower cost;
- easy onboarding;
- centralized migrations;
- simpler operations.
Trade-offs:
- logical isolation must be strong;
- noisy-neighbor risk exists;
- some enterprise customers may want stronger separation.
Bridge
Infrastructure is shared, but tenants receive stronger logical separation such as schema-per-tenant or database-per-tenant on shared infrastructure.
This can provide more isolation without operating a completely separate stack for every customer.
Silo
Selected tenants receive dedicated databases or infrastructure.
This gives strong blast-radius and performance isolation, but increases cost and operational work.
A real SaaS does not have to use one model for every customer. Pooling most tenants and isolating selected enterprise tenants is a valid architecture.
PostgreSQL: Do Not Trust Every Query to Remember tenant_id
A pooled database often starts with application filtering:
SELECT *
FROM orders
WHERE tenant_id = $1;
You should still write tenant-aware queries.
But relying only on every developer remembering the correct filter creates a fragile security boundary.
PostgreSQL Row-Level Security moves an additional isolation check into the database.
A simplified policy can look like:
ALTER TABLE orders ENABLE ROW LEVEL SECURITY;
ALTER TABLE orders FORCE ROW LEVEL SECURITY;
CREATE POLICY tenant_orders
ON orders
USING (
tenant_id = current_setting('app.tenant_id', true)::uuid
)
WITH CHECK (
tenant_id = current_setting('app.tenant_id', true)::uuid
);
Current PostgreSQL documentation defines USING as the rule for which existing rows can be accessed and WITH CHECK as the rule for rows created or modified.
RLS should be defense in depth, not an excuse to stop writing clear tenant-scoped application code.
The Database Role Matters
PostgreSQL documents two important RLS bypass cases:
- superusers and roles with BYPASSRLS always bypass row security;
- table owners normally bypass RLS unless FORCE ROW LEVEL SECURITY is used.
For normal SaaS traffic:
- do not connect as a superuser;
- do not grant BYPASSRLS;
- understand who owns tenant tables;
- use FORCE ROW LEVEL SECURITY where appropriate;
- test with the same role production uses.
An RLS policy tested only through an administrator connection can give false confidence.
Connection Pools Can Leak Tenant Context
Database pools reuse connections across requests.
This is dangerous if tenant context is stored in session state and is not safely replaced.
Conceptually:
Tenant A request
→ connection gets tenant A context
→ connection returns to pool
→ Tenant B receives same connection
Every transaction that touches tenant data should establish its trusted tenant context itself.
Use transaction-scoped context where your driver and pooling mode support it, and test the exact pool mode used in production.
Do not rely only on cleanup code that might be skipped after errors.
Tenant Isolation Belongs in the Schema Too
RLS is not the entire database design.
Tenant identity often needs to participate in:
- unique constraints;
- foreign-key relationships;
- indexes;
- frequent query predicates.
Suppose order numbers only need to be unique inside a tenant.
A global uniqueness rule is wrong.
The intended business rule is closer to:
UNIQUE(tenant_id, order_number)
The same applies to slugs, external IDs, category names, and other tenant-owned identifiers.
RLS protects access. Schema design protects data invariants.
Keep PostgreSQL Security Updates Current
PostgreSQL fixed CVE-2026-14666 on August 13, 2026.
The issue involved cached row-security policies remaining stale after certain role or ownership changes until cache invalidation or connection termination.
Fixed versions include PostgreSQL 18.6, 17.11, 16.15, 15.19, and 14.24.
The takeaway is simple: RLS is security infrastructure. Keep PostgreSQL patched and test permission-revocation behavior rather than treating RLS as a one-time configuration task.
Redis: Tenant Scope Is Part of the Cache Key
Redis can return data without touching PostgreSQL.
That means a bad cache key can bypass the isolation that RLS would otherwise provide.
Bad:
dashboard:user:42
Better when the data is tenant-specific:
tenant:t_91:dashboard:user:42:v3
Tenant scope belongs in cache keys for:
- cached entities;
- dashboard summaries;
- idempotency state;
- workflow state;
- tenant configuration;
- feature or entitlement snapshots;
- rate-limit state.
If authorization changes the representation, either include the relevant security dimension or cache below that transformation layer.
A fast cache hit is not useful if it returns another customer's data.
Security State Needs Stronger Cache Rules
Caching ordinary product data and caching permissions are not the same thing.
If membership is revoked but a permission cache lives for 30 minutes, access may continue for 30 minutes.
For security-sensitive state:
- use bounded TTLs;
- invalidate on membership and role changes;
- version important state where useful;
- define what happens if invalidation fails.
Cache performance must not silently redefine your access policy.
Queues: Tenant Context Must Survive the Async Boundary
Workers are a common place to lose tenant scope.
A job envelope can include:
job_id
tenant_id
actor_id
operation
resource_id
correlation_id
tenant_id should come from the trusted producer, not arbitrary user input.
The worker should:
- validate the job;
- establish tenant context;
- load the resource through a tenant-scoped path;
- execute;
- emit tenant-scoped logs, audit, and usage events.
Avoid workers that only load a record by resource_id if the resource is tenant-owned.
Use tenant_id plus resource identity, or let database-level isolation enforce the same boundary.
Shared Queues Also Need Performance Isolation
Tenant isolation is not only confidentiality.
One customer should not be able to flood a shared queue and delay every other customer.
Azure's multitenant messaging guidance explicitly calls out noisy-neighbor risk in shared messaging systems.
Useful controls include:
- per-tenant concurrency;
- per-tenant rate limits;
- maximum queued jobs;
- fair scheduling;
- separate worker pools for expensive workloads.
High-value or regulated customers can use dedicated queues when requirements justify the operational cost.
Object Storage Needs the Same Boundary
A common storage layout is:
tenants/{tenant_id}/documents/{file_id}
That is useful, but the path itself is not authorization.
The server should derive tenant_id, build the object path, authorize the resource, and issue narrow signed URLs for direct uploads or downloads.
Do not let the client freely choose another tenant prefix.
For stronger isolation, selected tenants can receive dedicated buckets, accounts, or storage resources.
Azure's multitenant storage guidance describes the same spectrum from shared paths and containers to dedicated storage accounts.
Search Is Another Data Store
Search often gets added after the primary database, which makes it an easy isolation blind spot.
Every indexed tenant document should preserve tenant attribution.
Every customer query should inject trusted tenant filtering on the server.
Options include:
- shared index plus mandatory tenant filter;
- logical namespace per tenant;
- dedicated index for high-assurance tenants.
The exact technology matters less than one rule:
> Tenant filtering is a security condition, not a frontend search option.
RAG and Vector Search Need Isolation Before Retrieval
RAG can turn a retrieval mistake directly into an AI data leak.
A safe path looks like:
authenticated request
↓
trusted tenant context
↓
embedding / search request
↓
mandatory tenant filter
↓
document authorization
↓
reranking
↓
LLM context
Do not retrieve globally and filter after the documents have already entered model context.
Depending on the vector database, use metadata filters, namespaces, native multi-tenancy, or dedicated collections for sensitive tenants.
A system prompt that says "only use this tenant's data" is not access control.
AI Tools Must Not Choose the Tenant
Imagine an agent tool:
update_order(tenant_id, order_id, status)
Allowing the model to provide tenant_id makes the security boundary part of model output.
A stronger design is:
trusted runtime already knows tenant
model provides order_id and status
server injects tenant context
tool authorizes and executes
The model proposes business arguments.
Trusted code controls identity, tenant, credentials, and authorization.
Prompt injection or hallucination must not be able to switch customer context.
Analytics Can Reintroduce Cross-Tenant Leaks
Operational events often flow into ClickHouse, warehouses, or reporting systems.
Preserve trusted tenant attribution through that pipeline.
Customer-facing analytics queries must always scope by tenant.
Internal company-wide analytics can intentionally aggregate tenants, but that should be limited to authorized internal use and should avoid unnecessary raw customer data.
An analytics database being internal does not make its API queries automatically safe.
Logs, Traces, and Audit Need Tenant Correlation
Tenant-aware observability helps diagnose both security and noisy-neighbor problems.
Useful fields include:
tenant correlation
service
operation
trace_id
duration
result
Do not dump customer payloads into logs just to make them tenant-aware.
Customer-facing audit trails should use a dedicated audit design with stronger integrity and retention rules.
Schedulers Need Current Tenant State
A scheduled job can run weeks after it was created.
At execution time:
- resolve the tenant;
- check relevant entitlement or feature state;
- verify the target still belongs to the tenant;
- apply current security policy;
- record the service or delegated actor.
Do not assume permission remains valid forever because the schedule was valid when created.
Noisy Neighbors Are Isolation Failures Too
Pooled infrastructure creates cost efficiency, but one tenant can consume disproportionate resources.
AWS and Azure both highlight this trade-off.
Useful controls include:
- API limits per tenant;
- database query budgets;
- queue concurrency;
- export/report limits;
- AI token or cost budgets;
- storage quotas;
- bulk-operation throttling.
Track enough per-tenant consumption to find abusive or simply very large workloads without creating uncontrolled metric cardinality.
Hybrid Isolation Is Often the Best Long-Term Shape
You do not need a separate stack for every customer.
A practical model can be:
Most tenants:
shared application
shared PostgreSQL
tenant_id + RLS
shared Redis
shared queues
shared storage with scoped paths
Enterprise tenant:
dedicated database
dedicated queue or worker pool
stronger storage boundary
stricter key management
Keep one higher-level application contract so moving a customer to stronger isolation does not create a second product.
Test Tenant Isolation Adversarially
Do not test only the happy path.
Database
Use Tenant A credentials to request a Tenant B resource.
Test SELECT, INSERT, UPDATE, and DELETE using the real production-style database role.
Cache
Create overlapping entity IDs in two tenants and prove keys cannot collide.
Queue
Feed a worker a resource belonging to another tenant and verify it refuses the operation.
Storage
Attempt to request or sign an object owned by another tenant.
Search and RAG
Create semantically similar documents in two tenants and prove retrieval never crosses the boundary.
AI tools
Prompt the model to request another tenant explicitly and confirm the server-side tenant context wins.
Support paths
Verify privileged access is explicit, authorized, and audited.
Tenant-isolation tests belong in CI, not only in an annual security review.
Common Failures
Trusting tenant_id from the client
Tenant scope must come from trusted identity.
Relying only on WHERE clauses
Use application scoping plus database enforcement such as RLS where appropriate.
Running the app as table owner or BYPASSRLS
That can bypass PostgreSQL row security.
Leaking session context through a connection pool
Set tenant context safely for each transaction.
Redis keys without tenant scope
A cache hit can become a cross-tenant leak.
Background jobs containing only resource_id
Workers still need tenant context.
Search filters controlled by the frontend
Tenant filtering is server-enforced.
RAG filtering after retrieval
Another tenant's document may already be in model context.
AI tools accepting arbitrary tenant IDs
The runtime should inject tenant context.
Unlimited shared workloads
A noisy tenant can still damage every other tenant's service.
A Practical Starting Architecture
For many SaaS products, a strong starting design is:
Identity:
tenant from authenticated membership
PostgreSQL:
shared schema
tenant_id on tenant-owned tables
RLS on tenant data
non-owner, NOBYPASSRLS app role
tenant-aware indexes and constraints
Redis:
tenant namespace in tenant-dependent keys
Queues:
trusted tenant_id in job context
worker re-establishes scope
per-tenant concurrency
Storage:
server-generated tenant paths
narrow signed access
Search / RAG:
mandatory tenant filters or namespaces
AI:
tenant injected by trusted runtime
Observability:
safe tenant correlation
per-tenant resource visibility
This provides strong logical isolation while keeping the system operationally simple.
Move selected tenants to dedicated resources only when compliance, performance, security, or commercial requirements justify it.
Production Checklist
- Tenant identity comes from trusted authentication or machine credentials.
- Tenant context crosses API, worker, scheduler, and tool boundaries.
- Every tenant-owned PostgreSQL table has an explicit isolation strategy.
- Production database roles cannot bypass intended RLS policies.
- Connection pooling cannot leak tenant state.
- Tenant-aware uniqueness and indexes are correct.
- Redis keys include tenant scope where results differ by tenant.
- Workers enforce tenant context.
- Object-storage access is server-scoped.
- Search and RAG enforce tenant filters before retrieval.
- AI tools cannot select their own tenant boundary.
- Analytics preserves tenant attribution.
- Shared resources have noisy-neighbor controls.
- Cross-tenant attack tests run automatically.
- Admin and support bypass paths are explicit and audited.
- PostgreSQL security updates are kept current.
Final Takeaway
The most dangerous SaaS architecture is one where every subsystem assumes another layer already handled isolation.
PostgreSQL RLS does not namespace Redis.
Redis does not protect a queue worker.
A queue worker does not protect object storage.
Object storage does not protect vector retrieval.
A vector filter does not protect an AI tool.
Strong multi-tenant SaaS architecture carries one trusted tenant context through every layer that can read, write, cache, retrieve, process, or expose customer data.
trusted tenant identity
↓
consistent tenant context
↓
isolation at every resource boundary
↓
tests proving cross-tenant access fails
You can still keep most infrastructure shared and cost-efficient.
The goal is not maximum separation everywhere.
The goal is to make sharing intentional and cross-tenant access impossible by design.
Sources and Further Reading
- AWS Prescriptive Guidance — Multi-tenant PostgreSQL: https://docs.aws.amazon.com/prescriptive-guidance/latest/saas-multitenant-managed-postgresql/introduction.html
- AWS Prescriptive Guidance — Row-Level Security: https://docs.aws.amazon.com/prescriptive-guidance/latest/saas-multitenant-managed-postgresql/rls.html
- AWS SaaS Lens — Tenant Isolation: https://docs.aws.amazon.com/wellarchitected/latest/saas-lens/tenant-isolation.html
- PostgreSQL 18 — Row Security Policies: https://www.postgresql.org/docs/18/ddl-rowsecurity.html
- PostgreSQL — CREATE POLICY: https://www.postgresql.org/docs/current/sql-createpolicy.html
- PostgreSQL Security — CVE-2026-14666: https://www.postgresql.org/support/security/CVE-2026-14666/
- Azure Architecture Center — Multitenant Storage and Data: https://learn.microsoft.com/en-us/azure/architecture/guide/multitenant/approaches/storage-data
- Azure Architecture Center — Multitenant Messaging: https://learn.microsoft.com/en-us/azure/architecture/guide/multitenant/approaches/messaging
- Azure Architecture Center — AI and Machine Learning in Multitenant Solutions: https://learn.microsoft.com/en-us/azure/architecture/guide/multitenant/approaches/ai-machine-learning

Discussion (0)