Audio work takes time. Separating stems, transcribing an hour of speech, running a multi-step pipeline over hundreds of files. None of it returns in a single request. So every audio API has to answer the same question: how does the caller find out when the work is done?
Most of the market answers "poll." We answer "we'll call you."
The problem with polling
Polling means your code submits a job, then loops: every few seconds, ask the server "done yet? done yet?" You keep asking until the answer is yes.
It works, but it's the wrong default:
- It's wasteful. The overwhelming majority of your status checks come back "not yet." You burn requests, and sometimes rate limit, asking a question whose answer is almost always no.
- It's laggy. Whatever your poll interval is, that's your worst-case delay after the job actually finishes. Poll every 10 seconds and you'll routinely learn about completions 10 seconds late.
- It's stateful in the wrong place. Your service has to remember every outstanding job and keep a loop alive for each one. That's bookkeeping you didn't ask for.
Completion webhooks flip it around
With a webhook, you hand us a URL when you submit the job. When the work finishes, we send a request to that URL. No loop, no polling, no guessing. Your code goes back to doing other things and gets woken up exactly once, when there's something real to react to.
On Scarleta this works for both single jobs (a recipe_completed event) and bulk jobs (a batch_completed event), so the same pattern covers one file or five hundred.
A couple of engineering details we care about, because they're the difference between a webhook you can build on and one you can't:
- It's SSRF-guarded. We won't deliver to private or loopback targets. A webhook endpoint has to be a real, external URL.
- It's idempotent, with retries. Delivery is attempted a few times with backoff, and a replay won't double-fire your handler. You can build on the assumption that "handled once" means once.
Why this is a wedge, not a nicety
"Push instead of poll" sounds like a small ergonomic. In practice it changes the shape of everything you build on top. Your integration stops being a polling daemon and becomes a plain request handler. Your latency stops being your poll interval and becomes network time. Your job bookkeeping moves off your plate.
That's the point of good infrastructure: it takes the tedious, easy-to-get-wrong part and makes it disappear. Don't poll. We'll call you when it's done.
Completion webhooks (single and batch), SSRF-guarding, and idempotent retry-with-backoff are live capabilities. See the Developers page. Note: webhook delivery is configured per request in the API payload; there is no self-serve dashboard for managing webhook endpoints today.