Skip to content

Queues and the scheduler

Background work in ulams runs on three queue connections, each with its own retry_after, and scheduled work runs from a long-lived scheduler loop. Both start by themselves in the api container; this page explains what they do so you can size, split and replicate them. The design is in ADR 0083 (queues) and ADR 0068 (scheduler lock).

retry_after is the time after which the queue hands a job that has not finished to another worker. If it is shorter than the job’s own timeout, a second worker runs the same job while the first is still working: a double model call, a double course clone. So every long job runs on a connection whose retry_after is above its timeout. The connections are defined in api/config/queue.php, each in a database and a redis variant that follows QUEUE_CONNECTION:

Connection Queue retry_after Runs
redis or database (QUEUE_CONNECTION) default (also broadcast, video) 90 s Ordinary jobs: e-mail, notifications, events
redis-builder or database-builder builder BUILDER_QUEUE_RETRY_AFTER, default 2400 s Course Builder runs and steps (1800 s timeout), Living Course source checks and analysis, the Adapt build
redis-long-job or database-long-job queue-long-job 19000 s Course clone and import, video processing (18000 s timeout), and tenant provisioning from the platform API

Other settings choose the connection and queue of one feature, for example COURSE_BUILDER_QUEUE_CONNECTION and COURSE_BUILDER_QUEUE, LIVING_COURSE_QUEUE_CONNECTION and LIVING_COURSE_QUEUE, ADAPT_QUEUE_CONNECTION and ADAPT_QUEUE, VIDEO_QUEUE_CONNECTION and VIDEO_QUEUE, LONG_JOB_QUEUE_CONNECTION and LONG_JOB_QUEUE. Leave them unset unless you deliberately route a feature elsewhere, and then keep the rule: worker --timeout at least the job timeout, and below retry_after. With QUEUE_CONNECTION=sync (tests) the jobs run inline and none of this applies. A test (QueueRetryAfterConfigTest) keeps the configuration consistent.

ULAMS_WORKERS_MODE chooses how the workers and the scheduler run:

Mode Where What runs
per-tenant default of the script, so production One long-lived process per tenant and queue kind, plus one scheduler loop per domain (below). Lowest latency, but each idle process holds a booted Laravel, about 140 MB: 29 processes with 7 domains
lean the default of api/docker-compose.yml, so local development and the demo profile; opt-in elsewhere No long-lived PHP process per tenant. See Lean mode

queue.sh (supervisor program laravel-multidomain-queue, switched off with DISABLE_QUEUE=true) runs workers.sh queue. It keeps one long-lived queue:work process per kind and per tenant, reads the list of tenants again every WORKERS_CHECK_INTERVAL seconds (default 10), so a new tenant gets workers without a restart, and starts a process again when it exits:

Worker Connection and queues --timeout Memory
default the default connection, queues default,broadcast,video 60 (Laravel default) 256 MB
builder <driver>-builder, queue builder 1800 512 MB
long <driver>-long-job, queue queue-long-job 18000 512 MB

Every worker stops after WORKERS_MAX_TIME seconds (default 3600) and is started again with fresh code. The platform itself gets a long-job worker (tenant provisioning) from workers.sh; its other queues are served by Horizon (DISABLE_HORIZON=true turns Horizon off), whose supervisor-builder (timeout 1800) and supervisor-long-job (timeout 18000) match the table.

The script reads these variables from the container’s own environment, not from the Laravel .env that LARAVEL_* variables are written to:

Variable Default Meaning
QUEUE_CONNECTION redis database or redis: which variant of the builder and long-job connections to use. Set it as a plain variable on the container if you run the database driver
COURSE_BUILDER_QUEUE_CONNECTION <driver>-builder Connection of the builder worker
COURSE_BUILDER_QUEUE builder Queue of the builder worker
LONG_JOB_QUEUE_CONNECTION <driver>-long-job Connection of the long-job worker. The same name config/queue.php reads; the old name LONG_JOB_CONNECTION is still accepted as a fallback
LONG_JOB_QUEUE queue-long-job Queue of the long-job worker
ULAMS_WORKERS_MODE per-tenant (lean in api/docker-compose.yml) lean or per-tenant, see above. Any other value stops workers.sh
ENABLE_HORIZON false In lean mode, start Horizon anyway
WORKERS_LEAN_INTERVAL 10 Lean mode: seconds between passes over the domains for the default queues
WORKERS_LEAN_HEAVY_INTERVAL 20 Lean mode: the same for the builder and long-job queues
WORKERS_LEAN_MAX_TIME 30 Lean mode: seconds a pass keeps taking new jobs of one domain and queue
WORKERS_CHECK_INTERVAL 10 Seconds between reads of the tenant list
WORKERS_MAX_TIME 3600 Seconds a worker or scheduler loop lives before it is restarted

The LARAVEL_-prefixed settings, BUILDER_QUEUE_RETRY_AFTER among them, go to Laravel as usual (see Environment).

  1. Run a worker pool on its own: start the api image with DISABLE_PHP_FPM=true, DISABLE_SCHEDULER=true and DISABLE_DB_MIGRATE=true (the worker role in High availability). Add replicas for more throughput; jobs are popped atomically, so several replicas are safe.
  2. After changing a queue setting, restart the workers on every tenant with php artisan queue:restart --domain=<host> (or restart the container).
  3. Jobs queued before the connections existed sit on the default queue and finish there.

With ULAMS_WORKERS_MODE=lean the same work runs from two supervisor programs that each hold one small shell loop, and PHP runs only while there is a pass to do:

  • workers.sh queue loops over the platform (unless MULTI_DOMAINS is set) and every tenant and runs php artisan ulams:tenant:work-once --queue=default,broadcast,video --timeout=60 for each, which is queue:work --stop-when-empty for that domain. It waits WORKERS_LEAN_INTERVAL seconds (default 10) between passes.
  • The same program starts one helper process for the builder and the long-job queues. It takes one domain at a time and runs work-once on <driver>-builder (queue builder, --timeout=1800) and then on <driver>-long-job (queue queue-long-job, --timeout=18000), the timeouts and connections of the per-tenant table above, so the retry_after rule of ADR 0083 holds. It waits WORKERS_LEAN_HEAVY_INTERVAL seconds (default 20) between passes. A domain is worked for at most WORKERS_LEAN_MAX_TIME seconds (default 30) before the next one, but a job that has started runs to its end: a long Course Builder step on one tenant delays the builder queue of the others.
  • workers.sh scheduler is one loop: at the start of each minute it runs ulams:tenant:schedule-loop --once --lock for the platform and every domain, so the hourly demo reset, the daily pruning and the Living Course polling run as before. --lock claims the minute with the lock described below, so a second container does not repeat a tick.
  • Horizon is not started (ENABLE_HORIZON=true starts it again; its workers then run next to the loop, which is harmless because jobs are popped atomically). Horizon only served the platform queues and its dashboard at /horizon.

Trade-offs: a job waits for the next pass, up to about the interval plus the length of a pass over the domains (a few seconds per domain), and every pass boots Laravel once per domain, which costs some CPU even when the queues are empty. Use per-tenant where latency matters, or lower the intervals. Code edits apply at the next pass without a restart.

scheduler.sh (supervisor program laravel-schedule, DISABLE_SCHEDULER=true turns it off) runs workers.sh scheduler: one ulams:tenant:schedule-loop per tenant and one for the platform. Each loop boots Laravel once, wakes at the start of every minute and runs that domain’s schedule:run in-process, then exits after WORKERS_MAX_TIME and is started again.

Run more than one scheduler container for failover and each loop would run every task twice, so the loop first claims the minute with an atomic lock in the cache store:

  • The lock is named ulams:schedule-tick:<tenant slug or platform>:<YmdHi> and lasts 90 seconds. The cache is Valkey with a prefix per tenant, so tenants do not share locks.
  • The replica that gets the lock runs the tick and does not release it, so a replica that is a few seconds late skips the minute. The others skip it as well.
  • A cache store without lock support (file, array) runs every tick, so use a shared store.
  • If the replica holding a minute dies mid-tick, that minute’s tasks are skipped; the next minute is taken by whoever gets there first. Keep the clocks of the replicas in sync (NTP): a clock a minute off can run a minute twice.
  • TENANCY_SCHEDULER_LOCK=false (default true) turns the lock off, for a deliberate second scheduler. php artisan ulams:tenant:schedule-loop --once --domain=<host> is a manual tick and never takes the lock.

The scheduler of the lean mode (one --once --lock tick per domain and minute) takes the same lock. A tick that lasts long delays the ticks of the domains after it, and one that slips into the next minute misses that minute’s tasks, so keep scheduled work short or queued (the demo reset runs in the background).

The hourly demo reset and the daily pruning jobs run from this scheduler, so a tenant whose scheduler is stopped stops resetting and pruning.