Just ask. Vocals gone, key found, transcript ready.
Upload your audio and tell Scarleta what you want in plain language. It figures out what to run, does the work, and shows you the results (players, charts, and files) right in the conversation. No code, no settings, no menus to learn.
How it works
No setup, no plugins. Just ask, and get the file back.
Upload your audio
Drop a track into the Scarleta chat: a song, a voice memo, a podcast, even a video.
Say what you want
Ask in plain language: "take the vocals out," "what key is this in," "transcribe this." The AI works out which tools to run, so you don't pick settings or write code.
Get results in the chat
Hear the processed audio in an inline player, see analysis as charts and tables, and download the files, all without leaving the conversation.
Built-in storage
Every result is saved to your library, 1 TB free.
The files you bring and the processed audio results you get are kept in one place, so you can come back, compare earlier versions, share a link, or ask across everything you’ve stored, whenever you want.
Recipes
Make the processed audio one step in a pipeline you run by name.
Chain it with other steps into a recipe you design once, then run on a single file in chat or on thousands at once.
See it in action
Just ask, and watch it happen
Simple, transparent pricing
Only pay for the audio you actually process.
1 token = 1 second of audio, minimum 1 token per job.
More than one trick
Scarleta does a whole lot more.
The same account handles all of it. Here are a few very different things you can do with your audio.
For developers
Ship the processed audio in your own product.
The same the processed audio is available over a clean REST API and an MCP server, priced by the second. Send a batch of files, get results back by webhook, and skip the polling.
Just ask and watch it get done
Start on your free tokens. Upload a track, tell Scarleta what you want, and get the results right in the chat. Top up only when you need more.
Frequently asked questions
What can I actually ask it to do?
The everyday audio jobs creators need: remove or isolate vocals and other parts, split a song into stems, find the key and chords, get the tempo, transcribe and caption speech, clean up noise, trim a clip, and more. You describe it and the AI works out how.
Do I need to know how to code?
No. That's the whole point. You type what you want in plain language and it does the work. There's nothing to install and no settings to configure.
How do I see the results?
Right in the chat. Processed audio plays in an inline player, analysis shows up as interactive charts and tables, and you can download the files it makes, like stems, a transcript, or the raw data as CSV or JSON.
How much does it cost?
You pay per second of audio, $0.10 per minute, from prepaid tokens (1 token = 1 second). New accounts start with 300 free tokens, and top-ups start at $10. No subscription.
Can developers do this in their own app?
Yes. The same conversational agent is available as an API you can build on. See the docs.