Skip to content
Audio Transcriber icon

Audio Transcriber

Speech-to-text with speaker labels

Audio Transcriber converts spoken audio into accurate text with speaker labels, word-level timestamps, and audio event tagging. It supports 90+ languages with automatic detection, making it suitable for meetings, interviews, podcasts, lectures, and any multi-speaker recordings.

Getting clean, structured transcripts normally requires a separate service and manual cleanup. This tool handles language detection, speaker diarization, timestamp granularity, and non-speech event tagging (laughter, applause, music) in a single call. The output includes both a full transcript string and a per-word array with timing and speaker IDs — ready for subtitles, summaries, or further analysis.

What you can do

  • Transcribe any audio file from a URL in 90+ languages with automatic language detection
  • Label different speakers separately (diarization) with or without knowing the speaker count
  • Get word-level or character-level timestamps for subtitle generation
  • Tag non-speech audio events like laughter, applause, and background music
  • Process MP3, WAV, M4A, FLAC, OGG, and other common formats
  • Turn a meeting recording into speaker-attributed notes, decisions, action items, risks, and optional PII-redacted audio

Who it's for

Podcast producers generating episode transcripts. Journalists and researchers transcribing interviews. Teams needing meeting notes with speaker attribution. Developers building transcription pipelines or subtitle generation workflows.

How to use it

  1. Call transcribe_audio with the URL of your audio file — language auto-detects if you don't specify
  2. Set diarize: true to get separate speaker labels; add num_speakers if you know the count for better accuracy
  3. Set timestamps_granularity: "word" if you need per-word timing for subtitle generation
  4. Enable tag_audio_events: true to capture laughter, applause, music, and other non-speech sounds

Getting started

For noisy recordings, run the audio through Audio Isolator first for cleaner transcription results. Then call transcribe_audio with the cleaned file URL.

Information

Price
From $0.005
Billing
1 free skill. The final price is shown before running. Failed paid calls do not charge.
ElevenLabs API Key
Optional: use your own ElevenLabs key instead of the platform default · Get key
AssemblyAI API Key
Optional key for meeting intelligence and regional processing · Get key
fal.ai API Key
Optional: use your own fal.ai key instead of the platform default · Get key
Prodia API Token
Optional: use your own Prodia token instead of the platform default · Get key
Higgsfield API Key
Optional: use your own Higgsfield key instead of the platform default · Get key
Photalabs API Key
Optional: use your own Photalabs key instead of the platform default · Get key
Google AI API Key
Optional: use your own Google AI key instead of the platform default · Get key
OpenRouter API Key
Optional: use your own OpenRouter key instead of the platform default · Get key
Runway API Secret
Optional: use your own Runway key instead of the platform default · Get key

Example workflows

Transcribe an interview with speaker labels

Turn a hosted audio recording into searchable text with optional speaker and word timing data.

  1. Pass the ToolRouter file ID or hosted audio URL to `transcribe_audio`.
  2. Enable diarization and word-level timestamps when you need speaker attribution or subtitle timing.
  3. Review the transcript, detected language, speaker labels, and timing data before using it downstream.

Guides using Audio Transcriber

Frequently Asked Questions

Can I transcribe an audio file to text?

Yes. Pass a ToolRouter file ID or hosted HTTP(S) URL to `transcribe_audio` to get the full transcript in one call.

Can it separate speakers and add timestamps?

Yes. Enable speaker diarization for labels, and request word-level timestamps when you need subtitle-style timing.

Which languages and audio formats work?

It supports 90+ languages with automatic detection and common formats including MP3, WAV, M4A, FLAC, and OGG.

Can it tag non-speech audio or handle a noisy recording?

Enable audio event tagging for laughter, applause, and music. For messy speech, run `audio-isolator` first.

Related Tools

Open Generate Video
Generate Video icon
Generate VideoTurn text or stills into video3 skills · from $0.005
Rated 3.0 of 5 —1
Open Generate Image
Generate Image icon
Generate ImageAI image generation, 20+ models6 skills · from $0.005
Rated 4.1 of 5 —7
Open Voice Generator
Voice Generator icon
Voice GeneratorText to speech with 1000+ voices3 skills · from $0.005
Rated 4.0 of 5 —1