Skip to content
Chif3n
4 min read

Notes from PyCon Cameroon 2026: Architecting Resilient Message Queues When the National Grid Blinks

When your server datacenter in Douala loses grid power twice a day, in-memory task queues disappear. Here is the Redis AOF persistence and Celery ACK acknowledgment pattern we presented at PyCon.

When European or North American developers talk about queue resilience, they plan for a 1-in-10,000 AWS zone outage or network partition.

When you run clinical dispatch queues in Cameroon, grid power doesn’t just fail as a black-swan event. It drops twice a day during regular load-shedding cycles or sudden equatorial downpours in Douala and Yaoundé.

If your message broker (RabbitMQ or Redis) buffers messages purely in volatile RAM, an ungraceful hardware cutoff destroys your queue:

  • A hospital blood request dispatched 8 seconds before the blackout disappears forever.
  • A critical emergency transfusion broadcast never reaches donors.
  • Or worse: on system reboot, an unacknowledged task executes twice, notifying the same donor twice and creating clinical confusion.

At PyCon Cameroon 2026, I gave a talk on architecting asynchronous task workers for environments with hostile electrical and network uptime. Here is the core pattern.


The Three Rules of Grid-Fault Tolerant Queues

1. Late Acknowledgment (task_acks_late = True)

By default, Celery marks a task as completed (ACK) the microsecond a worker picks it off the queue. If the machine dies 2 seconds into execution, the task is lost. With late ACKs, the message stays on the broker until the task successfully commits its changes to disk or returns.

2. Idempotency Keys with Redis Leases

Because a dying worker might reboot and re-run a partially executed job, every external side-effect (sending an SMS, debiting a hospital ticket, firing a WhatsApp message) must check an atomic idempotency lock:

# Acquire 10-minute unique idempotency key
if not redis.set(f"idemp:{ticket_id}", "locked", nx=True, ex=600):
    logger.warning("Duplicate execution trapped. Skipping side-effect.")
    return

3. Redis Append-Only File (appendfsync everysec)

Redis defaults to periodic RDB snapshots (e.g. every 5 minutes if 100 keys changed). A sudden power cut will lose up to 5 minutes of buffered donor messages. Configuring AOF with 1-second sync limits potential queue loss to a maximum of 1,000 milliseconds.


The Production Worker Codebase

Examine the production Celery configuration, idempotent task executor, Redis persistence parameters, and simulated drop test cases below. Click folders to open subdirectories, scroll the source code, copy snippets, or download files directly:

resilient-queue-pyconworkers/celery_app.py
"""
workers/celery_app.py
Celery application instance tuned for hard power-drop survival.
Presented at PyCon Cameroon 2026.
"""
 
from celery import Celery
import os
 
BROKER_URL = os.getenv("CELERY_BROKER_URL", "redis://localhost:6379/0")
BACKEND_URL = os.getenv("CELERY_RESULT_BACKEND", "redis://localhost:6379/1")
 
app = Celery("lifedrop_dispatch", broker=BROKER_URL, backend=BACKEND_URL)
 
app.conf.update(
# 1. Critical: Do NOT acknowledge tasks before execution completes
task_acks_late=True,
 
# 2. Re-queue tasks if a worker process crashes mid-execution
task_reject_on_worker_lost=True,
 
# 3. Only pre-fetch 1 message per worker process to minimize in-flight loss
worker_prefetch_multiplier=1,
 
# 4. Enforce strict serialization
task_serializer="json",
result_serializer="json",
accept_content=["json"],
timezone="Africa/Douala",
enable_utc=True,
 
# 5. Default retry policies for external network calls
task_annotations={
"workers.tasks.clinical_alerts.send_urgent_donor_alert": {
"rate_limit": "20/m",
"max_retries": 4,
"default_retry_delay": 15,
}
}
)
 

The Proof Bar: What Happened in Real Emergencies

When our primary staging server in Douala lost grid power for 42 minutes during the May 2026 storms:

  1. Zero Discarded Requests: 31 pending blood alerts stayed recorded in the local Redis AOF disk stream.
  2. Smooth Worker Recovery: Upon generator kick-in, Celery resumed the queue within 14 seconds without human intervention.
  3. No Duplicate SMS Floods: Every single emergency notification was sent exactly once to the designated donor candidate.
All writing

Comments

Comments coming soon. Set up Giscus on the repo to enable them.