1 TB free audio storage

Your words. Your labels. Scarleta scores the match.

Write the labels you care about in plain words: "crowd noise," "acoustic guitar," "angry tone," anything. Scarleta scores how well each one matches the audio. No fixed tag list to pick from, and nothing to train.

How it works

No setup, no plugins. Just ask, and get the file back.

Step 1

Upload your audio

Drop a clip into the Scarleta chat: a recording, a field capture, or a sample from your catalog. No plugins to install.

Step 2

Write your own labels

Give it the words you care about in plain language: "dog barking," "distorted," "calm voice." You choose the labels. There is no preset list to fit into and no model to train.

Step 3

Read the match scores

Get a score for how well each label matches the audio, laid out as a table or chart right in the conversation so you can rank, threshold, or triage.

Built-in storage

Every result is saved to your library, 1 TB free.

The files you bring and the label match scores results you get are kept in one place, so you can come back, compare earlier versions, share a link, or ask across everything you’ve stored, whenever you want.

1 TB free Share with a link Ask across your collection
Saved to your library1 TB free
Label match scores result
Ready to play, share, or reuse
Play Share Versions

Recipes

Make label match scores one step in a pipeline you run by name.

Chain it with other steps into a recipe you design once, then run on a single file in chat or on thousands at once.

Content moderation triage
Upload clipScore your labelsExport CSV
Catalog classifier
Batch via APIScore custom tagsFilter by threshold
Brand safety check
Upload recordingScore your labelsReview ranked results

See it in action

Score any audio against your own labels

Simple, transparent pricing

Only pay for the audio you actually process.

$0.10/ minute

1 token = 1 second of audio, minimum 1 token per job.

Prepaid, no subscription 300 free tokens to start Top up from $10

More than one trick

Scarleta does a whole lot more.

The same account handles all of it. Here are a few very different things you can do with your audio.

For developers

Do this at scale with the API

Running a whole library? Send a batch of files to a custom-label recipe and let a webhook call you back when each one is scored. Your labels are defined in the recipe.

curl https://api.scarleta.ai/v1/batch \
  -H "Authorization: Bearer $SCARLETA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "recipe_id": "rec_0001_custom_audio_labels",
    "audio_urls": [
      "https://example.com/clip.wav"
    ]
  }'

# 202 + batch_id. Poll GET /v1/batch/:id/results,
# or let the completion webhook call you when scoring is done.

Score your first clip on your free tokens

Upload a clip, write the labels you care about, and read the match scores. No fixed list and nothing to train. Top up only when you need more.

Frequently asked questions

What do I get back?

A match score for each label you wrote, showing how well it fits the audio. You can see the scores as a table or chart right in the chat, and export them as CSV or JSON for your own tools.

How is this different from 'identify the sounds'?

Sound identification tells you what is in a clip from open-ended detection: you do not pick the words. This is the opposite: you write your own labels and it scores the audio against exactly those. Bring your own vocabulary.

Do I have to train a model or pick from a preset list?

No. There is no training and no fixed list. You type whatever labels you want and it scores them straight away.

How much does it cost?

You pay per second of audio: $0.10 per minute, from prepaid tokens (1 token = 1 second). New accounts start with 300 free tokens, and top-ups start at $10. No subscription.

Can I run this on a lot of files from my own app?

Yes. Developers can send many files at once through our API and get called back by webhook when every one is done. See the docs.

Match Audio to Your Own Labels — Scarleta | Scarleta