Turn recordings into accurate, editable text.
Upload a meeting, interview, podcast, lecture, or video. CharaVox detects the language, separates speakers, and returns a searchable transcript with sentence-level timestamps.
Manual transcription is slow because every correction requires replaying the same audio. CharaVox Speech to Text creates a structured first draft with sentence timestamps, optional speaker labels, and direct playback so you can review the exact moment behind each line.
The tool accepts common audio and video formats, supports browser recording, and keeps long-running jobs recoverable after a page refresh. Export clean text for notes or download SRT captions for editing and publishing workflows.
Sign in to upload or record. This preview never sends audio.
Capabilities
Upload or Record
Drag in an existing audio or video file, choose a file from your device, or record speech directly from the microphone.
Language Detection
Let Fun-ASR detect the spoken language automatically or provide a language hint when you already know the recording language.
Speaker Separation
Enable diarization for meetings and interviews to group transcript segments by speaker, with an optional speaker-count hint.
Caption Exports
Copy the transcript or download TXT and SRT files generated from sentence-level timestamps.
Why creators choose it
Resume long transcriptions after refreshing or closing the page.
Review each sentence against its exact place in the source recording.
Separate speakers without manually labeling every turn in a conversation.
Pay only for effective speech duration at one credit per second.
Use cases
Frequently asked questions
CharaVox accepts common formats including MP3, WAV, M4A, MP4, WebM, OGG, FLAC, AAC, Opus, and AMR. The service verifies the real file size and media duration before creating a transcription job.
Yes. Turn on speaker separation for meetings, interviews, or multi-person recordings. You can optionally provide an expected speaker count from 2 to 100, while the transcript keeps speaker labels on sentence-level segments.
The tool reserves credits from the source duration before processing, then charges one credit per second of effective speech reported by Fun-ASR. Any unused reserved credits are released automatically when the job completes.
Yes. Completed jobs can be downloaded as plain TXT or SubRip SRT files. SRT exports use the sentence-level timeline returned by the transcription service.
The transcript remains until you delete it or close your account. Source files follow the plan retention period: 7 days on Free, 30 days on Creator, and 90 days on Pro.
Related AI tools
AI Voice Changer
Upload or record clean speech, transcribe it with Fun-ASR, and generate a private WAV in a different AI voice with CharaVox.
ExploreAI Voice Generator
Turn any script into natural AI speech in 50+ languages and 27,000+ community voices. Pick a voice, paste your text, and download studio-grade voiceover in seconds.
ExploreAI Storyteller
Turn any script into a multi-voice audio story. Assign a different voice to each character, add a narrator, and export a polished audio drama in minutes — no studio, no casting, no editing.
ExploreVoice Cloning
Upload a 30-second audio sample and create a custom AI voice model. Generate speech in 6 languages with license-safe watermarking and commercial-use rights.
ExploreTurn your idea into audio now.
Open the generator and ship downloadable audio in under a minute. No instruments, no DAW, no setup.