0083. Course Builder: a stage change is one transaction under the run lock
Generated from docs/decisions/0083-course-builder-stage-advance-under-the-run-lock.md
- Status: Proposed
- Date: 2026-10-09
Context and problem statement
Section titled “Context and problem statement”A generation run once stalled with several database queue workers. Reproduced with 2-4 queue:work
processes on the database connection and the fake LLM driver (ConcurrentWorkersTest): the run stayed
running at stage quizzes with every step done and an empty queue, or a stage was opened and closed
twice (two STEP_FINISHED events).
Cause: GenerationService::advance() closed a stage under the run lock but wrote the next stage as
<next>:opening and created the next stage’s steps afterwards, outside the lock, in openStage().
Between the two, a worker finishing its own step saw a stage with no steps, read it as “all done” and
advanced again. openStage() then wrote stage back to a value an overlapping advance had already
moved past. With no step left in flight nobody called advance() again, so the run never finished.
Not involved: job uniqueness (none), release/backoff (tries = 1), the SSE wake-up (a read-only cache
key). dispatchWindow() also counted running steps and claimed pending ones without a lock, so workers
could overfill the concurrency window.
Decision
Section titled “Decision”advance()closes the stage, writes the nextstageand creates the next stage’s steps in one transaction underlockForUpdateon the run. A stage without steps (no quizzes asked for) is closed in the same transaction. There is no:openingstate.- Events, progress and job dispatch follow the commit and belong only to the worker that made the transition. A late worker sees the next stage open with pending steps and does nothing.
dispatchWindow()counts and claims under the same run lock and dispatches after the commit.ConcurrentWorkersTestforks 1-4 realqueue:workprocesses on the database connection against one session and asserts the run finishes and each stage opens and closes once.
Consequences
Section titled “Consequences”- Workers briefly queue on the run row; the work inside the lock is a few queries.
- The separate
retry_afterhazard (below) is resolved by the amendment.
Amendment (2026-10-09): long jobs get their own queue connections
Section titled “Amendment (2026-10-09): long jobs get their own queue connections”The database and redis connections have retry_after = 90, while a Course Builder or Living Course step
runs up to 1800 s: a long real LLM call was handed to a second worker while the first still ran it (a double
call, double cost). The same held for other long jobs that ran on the default connection.
- New connections in
api/config/queue.php, each in adatabaseand aredisvariant, with their own queue name (a queue name is shared by all connections of one driver, so the worker that pops it decidesretry_after; the name must be dedicated):<driver>-builder, queuebuilder,retry_after2400 (BUILDER_QUEUE_RETRY_AFTER):RunJob,StepJob(1800 s, 1 try), Living CourseAnalyseGroupJob,ProgressRulesJob(1800 s),CheckSourceJob(900 s) and the AdaptBuildAdaptSource(now with$timeout = 900; the HTTP call to the builder stays at 300 s).<driver>-long-job, queuequeue-long-job,retry_after19000:ProcessVideoandCloneCourse(18000 s; the clone ran on the default connection so far).redis-long-jobexisted already.
- The jobs pick the variant that matches
QUEUE_CONNECTION(databaseorredis); any other default connection (sync) is used unchanged.COURSE_BUILDER_QUEUE(_CONNECTION),LIVING_COURSE_QUEUE(_CONNECTION),ADAPT_QUEUE(_CONNECTION),VIDEO_QUEUE(_CONNECTION)andLONG_JOB_QUEUE(_CONNECTION)still override. api/workers.sh queuestarts a third worker per tenant for the builder queue with--timeout=1800(the long-job worker keeps--timeout=18000); Horizon hassupervisor-builderwithtimeout1800. Rule: worker--timeout>= job$timeoutand <retry_after.- H5P and SCORM imports and PDF rendering are synchronous HTTP calls, not queued jobs, so
retry_afterdoes not apply to them. tests/Integrations/QueueRetryAfterConfigTest.phpasserts, per job and per driver, that its connection exists, has a dedicated queue and aretry_afterabove$timeoutplus a margin; that every job inpackages/*/src/Jobswith a timeout above the default connection’sretry_afteris routed; and that the worker and Horizon timeouts match.- Jobs already queued on the old connection (
defaultqueue) finish there; set the variables above, restart the workers (queue:restart) and run a worker forbuilder.
Amendment (2026-10-10): lean workers for local development and the demo profile
Section titled “Amendment (2026-10-10): lean workers for local development and the demo profile”With seven domains workers.sh started 29 long-lived PHP processes in the api container (a default,
builder and long-job queue:work per tenant, a scheduler loop per tenant, Horizon with its supervisors),
about 140 MB each and 4.1 GB at idle, in a Docker VM of 7.65 GB. Docker Desktop went down four times on
9 and 10 October 2026.
ULAMS_WORKERS_MODE=lean|per-tenant(default of the script:per-tenant;docker-compose.ymlsetslean, so local development and the demo profile; opt-in for production).- Lean:
workers.sh queueis one loop that runsulams:tenant:work-once(queue:work--stop-when-empty, ADR 0091) for the default queues of the platform and every tenant, plus one helper process that runs the builder and long-job queues, one domain at a time, with the per-tenant timeouts (--timeout=1800on<driver>-builder,--timeout=18000on<driver>-long-job). The worker timeout stays at least the job timeout and belowretry_after, so the rule of the first amendment holds in both modes.workers.sh scheduleris one loop that runsulams:tenant:schedule-loop --once --lockfor every domain at the start of each minute (the minute lock of ADR 0068 still applies). - Horizon is off in lean mode (
ENABLE_HORIZON=truebrings it back); it only served the platform queues and its dashboard. - php-fpm in development and demo uses
pm = ondemandwithPHP_FPM_MAX_CHILDREN(8); the demo profile lowers opcache to 192 MB, 16 MB interned strings and a 32 MB JIT buffer. - Consequences: a job waits for the next pass (about 10 s for the default queues, 20 s for builder and
long jobs, plus the pass over the domains); a running builder step on one tenant delays the builder
queue of the others; each pass boots Laravel once per domain.
tests/Integrations/WorkersModeTest.phpasserts the lean commands, connections and timeouts.