Give any AI a full audio toolbelt.
Connect your agent to Scarleta's MCP server with a key, and it can separate stems, transcribe speech, and detect key and tempo, as native tool calls, not hand-rolled REST plumbing. Billed per use, same prepaid tokens.
How it works
No setup, no plugins. Just ask, and get the file back.
Grab a key
Self-serve a customer key at signup. No sales call. The same key that authenticates the REST and batch API connects your AI to the MCP server.
Point your AI at the MCP server
Add Scarleta's MCP server to your agent's MCP config with your key. Claude, ChatGPT, or a custom agent, anything that speaks the Model Context Protocol.
Your AI calls the audio tools
The tools show up natively in your agent. It can separate stems, transcribe, or detect key and tempo on demand. You pay per second of audio it processes.
Built-in storage
Every result is saved to your library, 1 TB free.
The files you bring and the audio tool calls from your agent results you get are kept in one place, so you can come back, compare earlier versions, share a link, or ask across everything you’ve stored, whenever you want.
Recipes
Make audio tool calls from your agent one step in a pipeline you run by name.
Chain it with other steps into a recipe you design once, then run on a single file in chat or on thousands at once.
See it in action
Give any AI a full audio toolbelt via MCP.
Simple, transparent pricing
Only pay for the audio you actually process.
1 token = 1 second of audio, minimum 1 token per job.
More than one trick
Scarleta does a whole lot more.
The same account handles all of it. Here are a few very different things you can do with your audio.
For developers
Connect your agent
Add Scarleta's MCP server to your agent with your customer key, the same key and same per-second billing as the REST and batch API. The exact endpoint URL and the full tool list live in the developer docs.
{
"mcpServers": {
"scarleta": {
"url": "https://<see-developer-docs>/mcp",
"headers": {
"Authorization": "Bearer scar_live_your_key"
}
}
}
}Give your agent an audio toolbelt today
Grab a key and point your AI at the MCP server. It starts on your free tokens, billed per second of audio, no subscription. The full setup is in the developer docs.
Frequently asked questions
Which AIs can connect?
Any AI that speaks the Model Context Protocol: Claude, ChatGPT, or a custom agent you build. It connects to Scarleta's MCP server with your customer key.
What can my AI do once it's connected?
The same audio capabilities as the rest of the platform, exposed as native tool calls: separate stems, transcribe speech, detect key and tempo, and more. See the docs for the full tool list.
How is it billed?
Per use, on the same meter as everything else: 1 token = 1 second of audio, $0.10 per minute, from prepaid tokens. New accounts start with 300 free tokens, and top-ups begin at $10. No subscription.
Do I need a different key from the API?
No. Your customer key authenticates the MCP server the same way it authenticates the REST and batch API. Grab a key once and use it everywhere.
Is the MCP server actually live?
Yes. It is a live developer offering alongside the REST and batch API. Head to the developer docs for the connection details.