Oblive Docs
Developers

Worker Capacity

Share task workers in first-come, first-served order across Operator and departments.

workerReplicas defaults to four. Stack generation supplies the same value to the backend as WORKER_POOL_CAPACITY; Compose uses WORKER_REPLICAS for both backend capacity and worker replicas. Set the backend value to the actual shared task-worker capacity when deploying without Compose. Chat execution is separate from this task-worker pool.

Operator and department tasks share one first-come, first-served queue across organizations. The oldest eligible request starts first, ordered by availability time, creation time, and ID. Retries rejoin at their next availability time. Future, blocked, stale, cancelled, disabled, paused, or exhausted requests do not hold up eligible work. Task dependencies, human approvals, profile and department enablement, and run budgets still apply.

All worker slots can serve one department or Operator. There are no per-organization, per-profile, or background concurrency caps. The former maxConcurrentRuns fields are removed from stored organization/profile configuration and API contracts. Retained agent-pack catalogs can still be read; their obsolete cap is discarded during validation.

Tasks retain backend-owned executionClass metadata: standard or background. Organization discovery and automatic integration exploration are background work, and children inherit that ancestry. Neither this label nor stored priority changes FCFS admission. Existing runs are never preempted.

PostgreSQL owns queue order; Redis only delivers wakeups. Admission takes a global transaction advisory lock with a one-second wait bound, checks shared capacity and the oldest eligible request, and revalidates organization, profile, task, dependencies, and budgets under row locks. Out-of-order or duplicate wakeups cannot jump the queue. Deferral charges no attempt or run credit.

Admission creates the immutable attempt before preparing its execution context. Preparation failure records a retryable failed attempt and releases capacity through the existing bounded retry policy. If failure persistence is unavailable, the worker leaves delivery pending and lease recovery remains the durable fallback. Broken context preparation cannot keep an unclaimed request at the queue head indefinitely.

The scheduler’s due-retry phase runs before schedule materialization. Its internal cursor preserves PostgreSQL timestamp microseconds for the same FCFS ordering; it is separate from public list cursors. No agent-authored command changes. Agents keep using the existing task/run flow.

For an upgrade that must preserve active work, pause new work through organization work-dispatch control, wait for active runs, Chat responses, and executing actions to finish, then back up persistent state before restarting. Resume dispatch after migration and health checks. Pausing dispatch alone does not cancel a running attempt.

The PostgreSQL regression suite is apps/backend/tests/database/worker-fcfs.postgres.test.ts and runs in CI with TEST_DATABASE_URL. It covers ten slots serving one profile, concurrent claims, ordering, eligibility, retry timing, context failure, and migration preservation. The full context/dependency journey is also covered by apps/backend/tests/integration/worker-capacity.test.ts, enabled with RUN_DATABASE_INTEGRATION_TESTS=true and disposable local database/object-store configuration.