Your audio jobs run while you sleep. We call you when they're done.
18 pro audio tools, composed into named Recipes, delivered by webhook. Submit 500 files in one call, receive a clean completion event. Same engine over a batch REST API and a conversational Chat API.
The wedge
Don't poll. We'll call you.
Submit a batch and go do something else. The moment it finishes, Scarleta POSTs a batch_completed event to the URL you supply. No polling loop to babysit, no wasted requests.
- Fires on batch completion and on single-recipe runs.
- SSRF-guarded: private and loopback targets are refused, no POST sent.
- Three delivery attempts with backoff, then it stops. Idempotent, no double-POST on replay.
- The only secret is whatever you embed in your own webhook URL.
{
"event": ...,
"batch_id": ...,
"status": ...,
"total_items": ...,
"completed_items": ...,
"failed_items": ...,
"cancelled_items": ...,
"actual_total_cost": ...,
"completed_at": ...
}Exactly nine keys, in this order. See the reference for full types.
Four ways to build on Scarleta
The Batch API is the GA-validated flagship. Every surface hits the same audio engine. Pick the shape that fits how you build.
Bulk, async: the flagship
- Submit up to 500 files per batch; over the cap returns a coded 422.
- Per-tenant fair concurrency: one big batch can never starve another.
- Retry failed items by lineage, cancel mid-flight; in-flight items finish.
- Ledger-accurate cost and honest, coded errors on every failure.
POST /v1/batchGET /v1/batch/:idGET /v1/batch/:id/resultsConversational / agent
- Standard tier: submit a turn, then poll for the result.
- Streaming tier: stream the conversational response as it arrives.
- Metered per customer on the same per-second meter as everything else.
- Same engine as the web chat. You own the front end.
POST /v1/chat → GET /v1/chat/:job_idPOST /api/chat (streaming)An AI other AIs can use
- The full audio toolbelt, exposed over the Model Context Protocol.
- Any MCP-capable AI connects with a customer key.
- Separate stems, transcribe, detect key and tempo, and more. Natively.
- Billed per use, on the same prepaid per-second meter.
Composable pipelines
- Run a saved pipeline over audio; poll or receive a recipe-completed webhook.
- Create, save, and reuse multi-step pipelines instead of re-wiring calls.
- Discover a growing public library of recipes we curate (curation in progress).
- You only ever see the public library plus your own recipes and drafts.
POST /v1/recipe/execute → GET /v1/recipe/jobs/:job_idHonest, coded errors
Every batch failure returns a stable, documented code, never a vague 500. These are the nine you can handle for.
missing_audio_urlsempty / absent audio_urls
400missing_recipe_idno recipe_id
400invalid_jsonnon-JSON body
400invalid_max_retriesmax_retries at or below 0, or non-integer
400batch_too_largemore than 500 items
422batch_not_foundunknown id / cross-tenant read
404batch_already_finalcancel an already-terminal batch
409batch_paid_onlyfree-tier key
403workflow_binding_missingdeploy-regression guard
500Why teams switch to Scarleta
The parts that most audio APIs get wrong.
Webhooks, not polling
Don't poll. We call you the moment a job is done, on both batch and single-recipe runs.
Composable recipes
Design a multi-step pipeline once, then run it over audio. Your own reusable workflows.
Transparent per-second pricing
$0.10 per minute, 1 token = 1 second of audio, prepaid. No seats, no subscription.
Dual access: chat and REST
The same engine behind a conversational Chat API and a clean REST / batch API. Two front doors.
No double-charge for the same audio
An owner-scoped content-hash cache means you never pay twice to re-analyze the same file in your account.
Tenant isolation, by design
Another account's batch id returns a 404. Never a data leak. Cross-tenant reads are blocked.
The full audio tool suite
Grouped by what you get out. Call any of them directly, or chain them into a recipe.
Self-serve keys
Get a key. Start on the free grant.
Sign up and mint a key yourself, no sales call required. Every key is validated with HMAC-SHA256. Start free, then top up when you need scale.
scar_live_*Paid, productionscar_test_*Paid, stagingsk_free_*Free, one-time 300-token grantPricing
$0.10/ minute of audio
Prepaid token top-ups, not a subscription. 1 token = 1 second of audio. Retries never double-charge.
- Free tier: a one-time 300-token grant to try it.
- Minimum top-up is $10. No seats, no monthly plan.
- Saved cards, low-balance alerts, and auto-recharge available.
Get a key. Ship your first batch today.
Prepaid, per-second pricing. Start on the free grant, then top up when you need scale. No subscription, no seat count.