Queues and the scheduler
Background work in ulams runs on three queue connections, each with its own retry_after, and
scheduled work runs from a long-lived scheduler loop. Both start by themselves in the api
container; this page explains what they do so you can size, split and replicate them. The design
is in ADR 0083 (queues) and
ADR 0068 (scheduler lock).
Queue connections
Section titled “Queue connections”retry_after is the time after which the queue hands a job that has not finished to another
worker. If it is shorter than the job’s own timeout, a second worker runs the same job while the
first is still working: a double model call, a double course clone. So every long job runs on a
connection whose retry_after is above its timeout. The connections are defined in
api/config/queue.php, each
in a database and a redis variant that follows QUEUE_CONNECTION:
| Connection | Queue | retry_after |
Runs |
|---|---|---|---|
redis or database (QUEUE_CONNECTION) |
default (also broadcast, video) |
90 s | Ordinary jobs: e-mail, notifications, events |
redis-builder or database-builder |
builder |
BUILDER_QUEUE_RETRY_AFTER, default 2400 s |
Course Builder runs and steps (1800 s timeout), Living Course source checks and analysis, the Adapt build |
redis-long-job or database-long-job |
queue-long-job |
19000 s | Course clone and import, video processing (18000 s timeout), and tenant provisioning from the platform API |
Other settings choose the connection and queue of one feature, for example
COURSE_BUILDER_QUEUE_CONNECTION and COURSE_BUILDER_QUEUE, LIVING_COURSE_QUEUE_CONNECTION and
LIVING_COURSE_QUEUE, ADAPT_QUEUE_CONNECTION and ADAPT_QUEUE, VIDEO_QUEUE_CONNECTION and
VIDEO_QUEUE, LONG_JOB_QUEUE_CONNECTION and LONG_JOB_QUEUE. Leave them unset unless you
deliberately route a feature elsewhere, and then keep the rule: worker --timeout at least the
job timeout, and below retry_after. With QUEUE_CONNECTION=sync (tests) the jobs run inline and
none of this applies. A test (QueueRetryAfterConfigTest) keeps the configuration consistent.
Workers
Section titled “Workers”ULAMS_WORKERS_MODE chooses how the workers and the scheduler run:
| Mode | Where | What runs |
|---|---|---|
per-tenant |
default of the script, so production | One long-lived process per tenant and queue kind, plus one scheduler loop per domain (below). Lowest latency, but each idle process holds a booted Laravel, about 140 MB: 29 processes with 7 domains |
lean |
the default of api/docker-compose.yml, so local development and the demo profile; opt-in elsewhere |
No long-lived PHP process per tenant. See Lean mode |
Per-tenant mode
Section titled “Per-tenant mode”queue.sh (supervisor program laravel-multidomain-queue, switched off with DISABLE_QUEUE=true)
runs workers.sh queue. It keeps one long-lived queue:work process per kind and per tenant, reads
the list of tenants again every WORKERS_CHECK_INTERVAL seconds (default 10), so a new tenant gets
workers without a restart, and starts a process again when it exits:
| Worker | Connection and queues | --timeout |
Memory |
|---|---|---|---|
| default | the default connection, queues default,broadcast,video |
60 (Laravel default) | 256 MB |
| builder | <driver>-builder, queue builder |
1800 | 512 MB |
| long | <driver>-long-job, queue queue-long-job |
18000 | 512 MB |
Every worker stops after WORKERS_MAX_TIME seconds (default 3600) and is started again with fresh
code. The platform itself gets a long-job worker (tenant provisioning) from workers.sh; its other
queues are served by Horizon (DISABLE_HORIZON=true turns Horizon off), whose supervisor-builder
(timeout 1800) and supervisor-long-job (timeout 18000) match the table.
The script reads these variables from the container’s own environment, not from the Laravel .env
that LARAVEL_* variables are written to:
| Variable | Default | Meaning |
|---|---|---|
QUEUE_CONNECTION |
redis |
database or redis: which variant of the builder and long-job connections to use. Set it as a plain variable on the container if you run the database driver |
COURSE_BUILDER_QUEUE_CONNECTION |
<driver>-builder |
Connection of the builder worker |
COURSE_BUILDER_QUEUE |
builder |
Queue of the builder worker |
LONG_JOB_QUEUE_CONNECTION |
<driver>-long-job |
Connection of the long-job worker. The same name config/queue.php reads; the old name LONG_JOB_CONNECTION is still accepted as a fallback |
LONG_JOB_QUEUE |
queue-long-job |
Queue of the long-job worker |
ULAMS_WORKERS_MODE |
per-tenant (lean in api/docker-compose.yml) |
lean or per-tenant, see above. Any other value stops workers.sh |
ENABLE_HORIZON |
false |
In lean mode, start Horizon anyway |
WORKERS_LEAN_INTERVAL |
10 |
Lean mode: seconds between passes over the domains for the default queues |
WORKERS_LEAN_HEAVY_INTERVAL |
20 |
Lean mode: the same for the builder and long-job queues |
WORKERS_LEAN_MAX_TIME |
30 |
Lean mode: seconds a pass keeps taking new jobs of one domain and queue |
WORKERS_CHECK_INTERVAL |
10 |
Seconds between reads of the tenant list |
WORKERS_MAX_TIME |
3600 |
Seconds a worker or scheduler loop lives before it is restarted |
The LARAVEL_-prefixed settings, BUILDER_QUEUE_RETRY_AFTER among them, go to Laravel as usual
(see Environment).
- Run a worker pool on its own: start the
apiimage withDISABLE_PHP_FPM=true,DISABLE_SCHEDULER=trueandDISABLE_DB_MIGRATE=true(theworkerrole in High availability). Add replicas for more throughput; jobs are popped atomically, so several replicas are safe. - After changing a queue setting, restart the workers on every tenant with
php artisan queue:restart --domain=<host>(or restart the container). - Jobs queued before the connections existed sit on the
defaultqueue and finish there.
Lean mode
Section titled “Lean mode”With ULAMS_WORKERS_MODE=lean the same work runs from two supervisor programs that each hold one
small shell loop, and PHP runs only while there is a pass to do:
workers.sh queueloops over the platform (unlessMULTI_DOMAINSis set) and every tenant and runsphp artisan ulams:tenant:work-once --queue=default,broadcast,video --timeout=60for each, which isqueue:work --stop-when-emptyfor that domain. It waitsWORKERS_LEAN_INTERVALseconds (default 10) between passes.- The same program starts one helper process for the builder and the long-job queues. It takes
one domain at a time and runs
work-onceon<driver>-builder(queuebuilder,--timeout=1800) and then on<driver>-long-job(queuequeue-long-job,--timeout=18000), the timeouts and connections of the per-tenant table above, so theretry_afterrule of ADR 0083 holds. It waitsWORKERS_LEAN_HEAVY_INTERVALseconds (default 20) between passes. A domain is worked for at mostWORKERS_LEAN_MAX_TIMEseconds (default 30) before the next one, but a job that has started runs to its end: a long Course Builder step on one tenant delays the builder queue of the others. workers.sh scheduleris one loop: at the start of each minute it runsulams:tenant:schedule-loop --once --lockfor the platform and every domain, so the hourly demo reset, the daily pruning and the Living Course polling run as before.--lockclaims the minute with the lock described below, so a second container does not repeat a tick.- Horizon is not started (
ENABLE_HORIZON=truestarts it again; its workers then run next to the loop, which is harmless because jobs are popped atomically). Horizon only served the platform queues and its dashboard at/horizon.
Trade-offs: a job waits for the next pass, up to about the interval plus the length of a pass over
the domains (a few seconds per domain), and every pass boots Laravel once per domain, which costs
some CPU even when the queues are empty. Use per-tenant where latency matters, or lower the
intervals. Code edits apply at the next pass without a restart.
The scheduler and its minute lock
Section titled “The scheduler and its minute lock”scheduler.sh (supervisor program laravel-schedule, DISABLE_SCHEDULER=true turns it off) runs
workers.sh scheduler: one ulams:tenant:schedule-loop per tenant and one for the platform. Each
loop boots Laravel once, wakes at the start of every minute and runs that domain’s schedule:run
in-process, then exits after WORKERS_MAX_TIME and is started again.
Run more than one scheduler container for failover and each loop would run every task twice, so the loop first claims the minute with an atomic lock in the cache store:
- The lock is named
ulams:schedule-tick:<tenant slug or platform>:<YmdHi>and lasts 90 seconds. The cache is Valkey with a prefix per tenant, so tenants do not share locks. - The replica that gets the lock runs the tick and does not release it, so a replica that is a few seconds late skips the minute. The others skip it as well.
- A cache store without lock support (
file,array) runs every tick, so use a shared store. - If the replica holding a minute dies mid-tick, that minute’s tasks are skipped; the next minute is taken by whoever gets there first. Keep the clocks of the replicas in sync (NTP): a clock a minute off can run a minute twice.
TENANCY_SCHEDULER_LOCK=false(defaulttrue) turns the lock off, for a deliberate second scheduler.php artisan ulams:tenant:schedule-loop --once --domain=<host>is a manual tick and never takes the lock.
The scheduler of the lean mode (one --once --lock tick per domain and minute) takes the same lock.
A tick that lasts long delays the ticks of the domains after it, and one that slips into the next
minute misses that minute’s tasks, so keep scheduled work short or queued (the demo reset runs in the
background).
The hourly demo reset and the daily pruning jobs run from this scheduler, so a tenant whose scheduler is stopped stops resetting and pruning.