6.8 KiB
Initial abuse and capacity controls
Verdict
The initial production boundary requires hard safety controls before customer traffic. The current project lock and small in-memory trigger queue are useful local protections, but they do not provide multi-tenant admission, durable backpressure, storage safety, request isolation, bounded execution, or an operator stop mechanism.
The accepted contract is recorded in
ADR-0022 and
Spec.System.Capacity. Exact values
remain deliberately open until production-like calibration; the required
dimensions and failure behaviors do not.
Current implementation inventory
| Surface | Existing protection | Production gap |
|---|---|---|
| HTTP | Fastify's default 1 MiB body limit | No route-class, user, Org, IP, or administrative-write rate limits; request/connection/handler timeouts are unset |
| Project execution | Database lock permits one active run per Project | No platform or Org concurrency ceiling and no fair scheduler |
| Busy Project | Process-local FIFO, default five items and five-minute wait | Restart loses accepted work; expiry is logged and dropped without a durable state or user notification |
| Agent execution | Global max-turn default of 25 | No wall-clock, tool-call, output, memory, CPU, process, or Org-specific budget |
| Feishu messages | Message batching limits text count/characters | File/post attachment count and size are unbounded; downloads may buffer whole files |
| Storage | Workspace path boundary | No Project or Org byte quota, archive-expansion limit, or entity-count ceiling |
| Usage | Provider-reported token/cost fields and Org aggregation | Missing cost can disappear from totals; no soft-threshold notification or provider-mode attribution contract |
| Host | systemd lifecycle | No service CPU, memory, or task ceiling and no audited workload brake |
Most Organization member/team/project lists are also unpaginated, and usage aggregation can load all matching runs. Those query shapes must be bounded as part of enforcement rather than hidden behind larger process memory.
Accepted contract
Layered limits
- Every mandatory dimension has a versioned, non-overridable platform ceiling.
- An Organization OWNER/ADMIN may configure a lower policy limit.
- Effective limit is
min(platform, organization); absent Organization policy means the platform ceiling, never unlimited. - Missing or invalid mandatory production ceilings fail startup.
Admission and fairness
- Project exclusivity remains one active run.
- Platform and Organization concurrency limits apply before starting work.
- Available slots are scheduled fairly across Organizations, FIFO within an Organization, so one tenant cannot monopolize a global FIFO.
- Accepted work is durable and queryable. Queue length and wait are bounded.
- Overflow is explicit
capacity_exhausted(429plusRetry-Afterfor HTTP; a clear busy response for Feishu). Nothing is silently dropped. - Wait timeout produces terminal
EXPIRED, not a later automatic start. - The initiator may withdraw their queued request. Project
MANAGEor an Organization admin may cancel queued requests in their scope. Actor, reason, and notification are recorded.
Retry and idempotency
- Feishu
event_idand API idempotency identify one admission across retries. - Once a run starts, the whole run is never automatically replayed because tool side effects may already exist.
- Only bounded transport retries before model output or tool side effects are automatic. Later failure/timeouts require an explicit new request.
Request, file, storage, and entity safety
- Authenticated Web requests are limited by user, Organization, and endpoint class; signed Feishu events by Organization app, chat, and sender; administrative writes use stricter independent buckets.
- Admin surfaces require both a valid user session and the corresponding platform/Organization administrator authorization. OAuth callbacks and signed Feishu ingress remain necessary non-session entry points.
- Hard limits cover request body, single-file bytes, attachment count, archive expanded bytes/file count/depth, Project/Organization storage, and per-type Organization entity counts.
- Transfer is streamed and counted; partial files are removed after failure.
- Crossing a ceiling blocks new growth without auto-deleting or hiding existing resources. Read, export, delete, and administrator remediation stay usable.
Run budget
Every run is bounded by wall time, model turns, tool-call count, per-tool time,
output/log/event bytes, and process memory/CPU/child-process count. A breached
budget ends explicitly as TIMED_OUT or LIMIT_EXCEEDED, records the dimension
and observed usage, releases capacity, and never masquerades as success.
Usage and cost
- Token usage, provider-reported cost, run count, and duration are attributed by Organization, Project, Run, model, and Provider Connection.
- Missing provider cost is
unknown, never zero. - Thresholds are soft alerts in the pilot; they do not automatically stop runs and are not payment settlement.
- Org admins receive Organization alerts. Platform admins additionally receive anomalous-cost alerts for platform-managed connections; BYOK does not trigger platform financial alerts.
Emergency control
An audited Platform Administrator control supports Organization or global
DRAIN (reject new work, pause queue starts, finish active runs) and STOP_NOW
(also cancel active runs). Both require an actor, reason, timestamp, explicit
recovery, and user-visible outcomes. This is workload control, not Organization
deletion or account suspension.
Required production evidence
- Production configuration validation fails closed for every missing or invalid platform ceiling.
- Concurrent noisy-neighbor tests prove platform/Org/Project concurrency and Organization-fair scheduling without starvation.
- Restart and crash tests prove accepted queue entries, idempotency, expiry, cancellation, and notifications survive process loss.
- HTTP and Feishu load tests prove every rate-limit identity and typed overload response, including retry guidance.
- File/archive/storage probes prove streaming enforcement, cleanup, and continued read/delete access while over limit.
- Hung provider/tool and resource-exhaustion probes prove every run/process budget terminates, records cause, and releases locks/capacity.
- Usage tests prove provider-mode attribution, unknown-cost handling, alert routing, and absence of hard monetary blocking.
DRAINandSTOP_NOWdrills prove scoped, audited, reversible behavior.- Production-like calibration publishes the tested numerical defaults and capacity headroom used by the release gate.