Every word, every speaker, timed to the second.
Upload an interview, podcast, or voice memo and ask for the transcript. Scarleta writes out what is said, time-stamps it for captions, and labels who spoke, right in the chat.
How it works
No setup, no plugins. Just ask, and get the file back.
Upload your audio
Drop an audio file into the Scarleta chat: an interview, a podcast episode, a recorded meeting, or a voice memo. Nothing to install.
Ask for the transcript
Type it in plain language: "transcribe this and label who's speaking." The AI figures out which analysis to run, with no settings and no code.
Get text, captions, and speakers
Read the full transcript with word-level timestamps for captions, and see who-said-what speaker labels right in the conversation. Export the data when you need it.
Built-in storage
Every result is saved to your library, 1 TB free.
The files you bring and the clean transcript results you get are kept in one place, so you can come back, compare earlier versions, share a link, or ask across everything you’ve stored, whenever you want.
Recipes
Make a clean transcript one step in a pipeline you run by name.
Chain it with other steps into a recipe you design once, then run on a single file in chat or on thousands at once.
See it in action
Transcript, captions, and who spoke, automatically
Simple, transparent pricing
Only pay for the audio you actually process.
1 token = 1 second of audio, minimum 1 token per job.
More than one trick
Scarleta does a whole lot more.
The same account handles all of it. Here are a few very different things you can do with your audio.
For developers
Do this at scale with the API
Running a whole archive? Send a batch of files to a transcription recipe and let a webhook call you back when each transcript, its captions, and its speaker labels are ready.
curl https://api.scarleta.ai/v1/batch \
-H "Authorization: Bearer $SCARLETA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"recipe_id": "rec_0002_transcribe_caption_diarize",
"audio_urls": [
"https://example.com/interview.mp3"
]
}'Get your transcript in minutes
Start on your free tokens: upload an interview, ask for the transcript, and get text, captions, and speaker labels back in the chat. Top up only when you need more.
Frequently asked questions
What do I get back?
A transcript of what is said, word-level timestamps you can use for captions, and speaker labels showing who spoke. You can export the transcript and analysis as JSON or CSV to use in your own tools.
Can it tell who's speaking?
Yes. It labels who-said-what (speaker diarization), so a multi-person interview or panel comes back separated by speaker.
Can it translate to English?
Yes. Alongside the transcript, Scarleta can translate the speech to English.
How much does it cost?
You pay per second of audio: $0.10 per minute, from prepaid tokens (1 token = 1 second). New accounts start with 300 free tokens, and top-ups start at $10. No subscription.
Do I need to install anything?
No. Upload your audio in the Scarleta chat and ask. It runs in your browser with nothing to download.
Can I transcribe a lot of files from my own app?
Yes. Developers can send many files at once through our API and get called back by webhook when every one is done. See the docs.