Batch Calls Without Tripping the Rate Limit
Send fifty requests to a model provider in a tight loop and around the tenth you get a 429. Implement two functions: one retries a single call with capped, jittered backoff, the other runs a batch through it.
ask_with_backoff(llm, prompt, max_retries=5, base_delay=1.0, max_delay=30.0, jitter=None, sleep=time.sleep)
- On a
RateLimitErrororOverloadedError, wait and try again. Wait exponentially longer each time:base_delay, then double it, then double again. - Cap the wait at
max_delay. No single wait may be longer than the cap, however many retries came before it. - Jitter.
jitter, when given, is a function returning a number in[0, 1). Call it once per wait and waitdelay * (0.5 + 0.5 * that number), wheredelayis the capped delay, so each wait lands between half the delay and all of it. Withjitter=None, wait the delay itself. - Any other error is your request being wrong, not the provider being busy. Re-raise it immediately.
- After
max_retriesretries, give up and raise the last error. sleepandjitterare injectable. Tests pass a recorder and a fixed sequence to check what you waited without waiting.
ask_all(llm, prompts, **kwargs) runs every prompt through the first function and returns the replies in order. Two kinds of failure, two different answers:
- A prompt whose request is wrong (any
APIErrorthat is not one of the two busy signals) fails alone: putNonein its place and carry on. - A prompt that runs out of retries means the provider is still down: let that error out and stop.
The model is scripted to fail on cue. Your loop is the only thing under test.