• BullMQ
  • Node.js
  • LLM APIs

Do Not Retry an LLM Rate Limit, Reschedule It

The short version

A rate limit is not a failure, so do not spend retry attempts on it. Work out which limit you hit, then delay the job until that limit resets. Retrying a daily quota error immediately only fills your dead letter queue while hours of quota remain.

Two very different 429 errors

LLM providers enforce several limits at once. Requests per minute and tokens per minute reset within a minute. A daily quota resets once a day. Both can come back as the same HTTP status, 429, and that is where the trouble starts.

If a job hits the per minute limit, waiting 30 seconds fixes it. If it hits the daily quota, waiting 30 seconds fixes nothing, and neither does waiting an hour. The correct delay depends on which limit tripped, and the status code alone does not tell you. Read the error body, since most providers say which quota was exceeded, and check your provider’s documentation for when each one resets.

What the default retry does to you

A queue with five attempts and exponential backoff is a reasonable default for network blips. Pointed at a daily quota it is a disaster. In LeadForge AI the default behaviour burned all five retries within about thirty seconds and then dead lettered the job, while hours of quota window still lay ahead. The job was fine. The system had simply given up on it far too early.

The fix was a change in how I think about it. A rate limit says “not now”, and a failure says “this did not work”. Only the second deserves a retry attempt.

Rescheduling in BullMQ

BullMQ has a mechanism for exactly this. Inside the worker you call rateLimit with a delay and then throw RateLimitError. The job goes back to waiting without being counted as a failed attempt.

JavaScript
import { Worker } from 'bullmq'

const worker = new Worker(
  'enrich-leads',
  async (job) => {
    try {
      return await callModel(job.data)
    } catch (error) {
      const waitMs = cooldownFor(error)
      if (waitMs) {
        await worker.rateLimit(waitMs)
        throw Worker.RateLimitError()
      }
      throw error // real failures still use normal retries
    }
  },
  { connection }
)

Notice that rateLimit applies to the worker’s queue, so every job on it waits, not just the one that failed. That is what you want when the quota is shared by all jobs. If one job is the problem, look at moving that single job to the delayed set instead.

The function that matters is cooldownFor. It reads the error and returns how long to wait, or nothing if the error is not a rate limit at all.

JavaScript
function cooldownFor(error) {
  if (error.status !== 429) return 0

  // Honour the server when it tells you how long to wait.
  const retryAfter = Number(error.headers?.['retry-after'])
  if (retryAfter) return retryAfter * 1000

  // Daily quota: wait until the provider's reset, not a few seconds.
  if (/per day|daily/i.test(error.message)) return msUntilQuotaReset()

  // Per minute limit: a minute is enough.
  return 60_000
}

The message patterns above are examples. Log a few real errors from your provider and match on what they actually say.

If you have more than one API key

NoteMind uses a pool of Gemini keys, round robin, with a cooldown per key that respects the difference between daily and per minute limits. A key that hit its per minute limit should come back soon. A key that hit its daily limit should sit out until the reset. Here is a minimal version of that idea.

JavaScript
export class KeyPool {
  constructor(keys) {
    this.slots = keys.map((key) => ({ key, availableAt: 0 }))
    this.cursor = 0
  }

  next(now = Date.now()) {
    for (let i = 0; i < this.slots.length; i++) {
      const index = (this.cursor + i) % this.slots.length
      const slot = this.slots[index]
      if (slot.availableAt <= now) {
        this.cursor = (index + 1) % this.slots.length
        return slot
      }
    }
    return null // every key is cooling down
  }

  cool(slot, ms) {
    slot.availableAt = Date.now() + ms
  }

  earliestAvailable() {
    return Math.min(...this.slots.map((slot) => slot.availableAt))
  }
}

When next returns null you do not fail anything. You pass the gap until earliestAvailable to the queue and reschedule, exactly as before.

A short checklist

  • Rate limit errors never consume a retry attempt.
  • The delay comes from the error (retry after header, or which quota was hit), not from a constant.
  • Real failures, like bad input or a crashed parser, still retry a few times and then land in the dead letter queue.
  • Every dead lettered job has a reason you can read, so you can tell “gave up” from “was never going to work”.

LeadForge AI runs verification, scoring and outreach through BullMQ workers on Redis, and the full design is in the LeadForge AI case study. NoteMind’s key pool is covered in the NoteMind case study.

From the projectLeadForge AIFind the right businesses. Understand them. Start the right conversation.Read the case study →

Keep reading