Skip to content

From the team

Running Laravel queues in production without losing jobs

Queues work on a laptop. In production, workers run for days, jobs fail halfway and deploys land mid-batch. Here is what to decide deliberately.

5 min read

Laravel queues are the part of an application that runs after the HTTP response has already been sent: sending the email, generating the PDF, syncing the order to the accounting system. This article is for developers and technical leads who have queues working locally and now need them to behave in production, where workers run for days, jobs fail halfway through and a deploy lands in the middle of a batch. None of it is exotic. Most of it is deciding a few things deliberately instead of accepting defaults.

What should go in a queued job?

The test is simple: if the user does not need the result to render the next screen, it belongs in a job. Emails, notifications, webhooks to third parties, image processing, report generation, anything that talks to a network you do not control. The request should write to the database, dispatch, and return.

Two rules keep jobs sane. First, pass identifiers, not objects. SerializesModels will happily serialize a model, but the job then runs against the state of that model at dispatch time in your head and at run time in reality. Dispatch with an ID and load fresh data inside handle. Second, keep each job doing one thing. A job that sends the invoice and updates the ledger and notifies the account manager will fail one-third of the way through and retry all three. Split it, or chain it with Bus::chain.

Redis or database driver for Laravel queues?

The database driver is fine for getting started and for low-volume internal tools. It needs no extra infrastructure, and a failed job is a row you can inspect with a SQL client. Its weakness is contention: every worker polls the same table, and under load you see lock waits and a jobs table that grows faster than it is pruned.

For anything customer-facing, use redis. It is faster, it handles many workers without lock contention, and it is the driver Laravel Horizon is built for. If you are already running Redis for cache or sessions, the cost of adding a queue connection is close to nothing. Keep the queue on a separate Redis database or instance from the cache, so that a cache:clear or an eviction policy configured for cache never touches pending jobs.

Horizon earns its place the day you have more than one queue. It gives you a dashboard of throughput, wait times and failures per queue, a supervisor configuration in code rather than in a systemd file, and horizon:terminate for graceful restarts. Define queues by priority, not by feature: high for anything a user is waiting on, default for the rest, low for reports and exports. Give high its own worker pool so a slow export can never delay a password reset.

How do retries, backoff and idempotency work together?

Set $tries and $backoff on every job explicitly. A webhook to an external API might use three tries with a backoff of [10, 60, 300] seconds; a job that reads from your own database probably wants one try, because if it fails once it will fail again and retrying only hides the bug.

Retries only help if a job can run twice safely. That is idempotency, and it is the part teams skip. Before sending an email, check whether it has been sent. Before charging a card, use the payment provider's idempotency key. Before inserting, use updateOrCreate or a unique constraint and catch the violation. Write each job as if it will be run again after a crash at any line, because it will be.

Use ShouldBeUnique for jobs that should not be queued twice concurrently for the same entity, and the WithoutOverlapping middleware for jobs that may be queued but must not run in parallel. Set a timeout lower than the worker's retry_after value, otherwise a slow job is handed to a second worker while the first is still running it, which is the most common way to get duplicate side effects.

How do you monitor failed jobs and restart workers on deploy?

A failed job is not an error message; it is a row in failed_jobs waiting for someone to notice. Wire Queue::failing or the JobFailed event to whatever the team already looks at, and alert on the size of the failed table rather than on individual failures. Retry with queue:retry after fixing the cause, not before, and prune the table on a schedule with queue:prune-failed.

Watch the age of the oldest pending job per queue and the number of workers actually alive; Horizon shows both, and without it queue:monitor can alert on queue size.

Workers load code once and keep it in memory. After a deploy, every running worker is still executing the previous release until it is told to stop. Make queue:restart (or horizon:terminate) the last step of every deploy script, and run workers under a process manager such as Supervisor or systemd so they come back up. The restart is graceful: each worker finishes its current job first.

Finally, treat queue configuration as part of the delivery, not a post-launch chore. On our own projects it sits in the fixed-price scope from the first phase, because a queue set up in a hurry after launch is where the quiet data loss happens.