forked from EduCraft/curriculum-project-hub
docs: define SaaS capacity controls
This commit is contained in:
@@ -0,0 +1,88 @@
|
||||
# ADR 0022: Layered Capacity Admission And Workload Safety
|
||||
|
||||
## Status
|
||||
|
||||
Accepted.
|
||||
|
||||
## Context
|
||||
|
||||
The initial SaaS deployment shares one production host across Organizations.
|
||||
The current Hub only has a one-run-per-project lock, a process-local five-item
|
||||
trigger queue, a five-minute queue wait, a global max-turn setting, and
|
||||
Fastify's default request-body limit. It has no global or Organization agent
|
||||
capacity, durable backpressure, fair scheduling, storage quotas, run wall-clock
|
||||
budget, multi-dimensional rate limiting, or emergency workload control.
|
||||
|
||||
Leaving these limits implicit lets one Organization monopolize the host, lets
|
||||
accepted work disappear on restart, and makes overload look like success. A
|
||||
single global FIFO would preserve arrival order while still permitting one
|
||||
tenant to occupy every execution slot. Unlimited queues or automatic whole-run
|
||||
replay would hide overload and can duplicate tool side effects.
|
||||
|
||||
## Decision
|
||||
|
||||
Every mandatory capacity dimension has a non-optional **Platform Capacity
|
||||
Ceiling**. An Organization OWNER/ADMIN may configure a lower **Organization
|
||||
Policy Limit**; the **Effective Limit** is the lower value, or the platform
|
||||
ceiling when the Organization value is absent. There is no unlimited fallback.
|
||||
Production refuses to start when a mandatory platform ceiling is absent or
|
||||
invalid. Platform ceilings are changed through versioned deployment
|
||||
configuration in the initial release; live platform-admin editing is deferred
|
||||
until audited approval and rollback exist.
|
||||
|
||||
Agent work uses a durable, bounded admission queue. Scheduling is fair across
|
||||
Organizations and FIFO within an Organization, while the existing one-active-
|
||||
run-per-Project rule remains. Capacity saturation queues only within effective
|
||||
length and wait limits. Queue overflow is an explicit `capacity_exhausted`
|
||||
response with retry guidance; an accepted request that waits too long becomes
|
||||
`EXPIRED`, is never started later, and notifies its initiator. Initiators may
|
||||
withdraw their own queued requests; Project `MANAGE` and Organization admins
|
||||
may cancel queued requests in their scope. Every transition is auditable.
|
||||
|
||||
The platform never automatically replays a whole Agent Run after execution has
|
||||
started. Limited transport retries are allowed only before model output or tool
|
||||
side effects. Otherwise the run ends explicitly and a retry is a new user-
|
||||
initiated request. Feishu event identity and API idempotency prevent delivery
|
||||
retries from creating duplicate admissions.
|
||||
|
||||
The first production gate requires hard limits for request rate/body size,
|
||||
agent concurrency, queue length/wait, files and attachment counts, archive
|
||||
expansion, Project and Organization storage, Organization entity counts, run
|
||||
wall time/turns/tool calls/output, individual tool duration, and execution
|
||||
process memory/CPU/process count. Reaching an entity or storage ceiling blocks
|
||||
new growth but does not delete, hide, or lock existing data; read, export,
|
||||
delete, and administrator remediation remain available.
|
||||
|
||||
Admin surfaces require an authenticated user plus the corresponding platform
|
||||
or Organization administrator authorization. Authenticated Web traffic is
|
||||
limited by user, Organization, and endpoint class; signed Feishu traffic is
|
||||
limited by Organization application, chat, and sender; administrative writes
|
||||
use separate stricter buckets. Rejections and truncation are never silent.
|
||||
|
||||
Token, provider-reported cost, run count, and duration are attributed by
|
||||
Organization, Project, Run, model, and Provider Connection. They are statistics
|
||||
and soft alerts in the initial release, not hard spend controls or settlement.
|
||||
Missing provider cost remains unknown rather than zero. Organization admins
|
||||
receive their alerts; platform admins additionally receive anomalous-cost
|
||||
alerts for platform-managed Provider Connections, not BYOK financial alerts.
|
||||
|
||||
Platform administrators have audited, reason-bearing, reversible workload
|
||||
brakes at Organization and platform scope. `DRAIN` rejects new work and pauses
|
||||
queue starts while allowing active runs to finish. `STOP_NOW` also cancels
|
||||
active runs. These controls do not delete or suspend the Organization.
|
||||
|
||||
Exact numerical ceilings remain open until production-like capacity tests
|
||||
measure the supported host and provider behavior. Defaults must be checked in,
|
||||
validated, observable, and proven at the release gate rather than guessed as
|
||||
domain constants.
|
||||
|
||||
## Consequences
|
||||
|
||||
- The current in-memory `TriggerQueue` and global-only agent settings are
|
||||
implementation divergences, not acceptable production fallbacks.
|
||||
- Durable ingress, run/lock recovery, authorization, telemetry, and capacity
|
||||
enforcement must share one Organization-scoped admission identity.
|
||||
- HTTP overload uses typed `429` responses with `Retry-After`; Feishu reports a
|
||||
clear busy/expired/canceled outcome to the initiating user.
|
||||
- Capacity calibration is a release prerequisite and must include noisy-neighbor,
|
||||
restart, queue expiry, file/archive, provider-stall, and workload-brake tests.
|
||||
Reference in New Issue
Block a user