🎧

Best Speech-to-Text Apps for Meetings and Recordings

Meeyra Team17 min read0September 9, 2026

The best speech-to-text app for meetings is the one that matches your audio, not the one with the highest advertised accuracy. Cloud AI transcription turns an hour of recorded meeting audio into text in roughly 3–10 minutes for $0.05–$0.25 per audio minute, at a word error rate of about 8–12% on real multi-speaker recordings.

That last number is the one nobody puts on a pricing page. Vendors advertise "99% accuracy," and they are not lying — they are quoting clean, single-speaker, studio-grade audio. Your recording is four people on laptop microphones, one of them in a car. This guide covers what actually changes the result: the five ways to convert a recording to text, how accuracy differs by language, which file formats work, what it costs, and the consent question you need to settle before you upload anything.

Table of Contents

Live Transcription vs Recorded File Transcription

Before comparing apps, split the problem in two. These are different jobs with different tools.

Live transcription happens while people are talking. The text appears during the call, latency matters more than perfection, and the output doubles as accessibility support for anyone who reads faster than they listen. If this is what you need, the decision belongs to your meeting platform rather than a separate app — see how to turn on live captions in a video call for the platform-by-platform settings.

Recorded file transcription happens afterward. You have an MP3, an M4A from a phone recorder, or an MP4 export from a call, and you need searchable text. Here latency does not matter at all — a model can take four passes over the audio, use context from later in the file to correct earlier words, and produce a cleaner result than any live system.

This article is mostly about the second job. The distinction matters commercially too: if your meetings already produce a transcript automatically, you never buy a transcription tool at all. That is why the strongest cost saving in this category is usually not a cheaper per-minute rate — it is not needing the second tool.

Five Ways to Turn a Recording Into Text

Almost every product on the market is a variation of one of these five approaches. Pick the row that fits your situation, then shop inside it.

ApproachAccuracy on real meeting audioLanguage coverageCostWhere your audio goesBest for
Office suite built-in transcriptionModerate; struggles with multiple speakersBroad for major languages, thin elsewhereIncluded with an existing subscription, capped monthly minutesVendor cloudOccasional single-speaker files when you already pay for the suite
Cloud AI transcription serviceGood; speaker labels included40–100 languages, quality varies by tier$0.05–$0.25 per audio minute, or a monthly planVendor cloudRegular volume, interviews, podcasts, research
Open-source model run locallyGood, and tunable; slow without a GPUAround 100 languages, uneven below the top tierFree software, your own hardware and timeNowhere — stays on your machineConfidential files, technical teams, offline work
Professional human transcriptionHighest; 99%+ on decent audioDepends on the agency's linguists$1–$3 per audio minute, higher for rush jobsAgency staff and their systemsLegal, medical and any transcript that will be quoted
Meeting platform with built-in transcriptGood, and captured at the sourceDepends on the platformIncluded in the planYour meeting providerTeams who want the transcript to exist without a second step

The last row deserves a note, because it is the one people forget. Meeyra generates a timestamped, translated transcript from the meeting itself, in 42+ languages, so the audio never becomes a file you have to process later. The live translation feature writes each speaker's words in every participant's language while the call is happening.

How Accurate Is AI Transcription in Real Meetings

Word error rate (WER) is the standard measure: the percentage of words the system gets wrong through substitution, deletion or insertion. Lower is better, and context is everything.

On clean read-aloud English — the LibriSpeech benchmark — leading models sit near 2–3% WER. OpenAI's Whisper research paper reports results in that range for its largest model, which is where the marketing figure of "97–99% accuracy" comes from.

Real meetings are harsher. Published benchmarks and vendor-independent testing put the same class of model at roughly 8–12% WER on multi-speaker conference audio, rising toward 15–25% on noisy phone-quality recordings. Strong accents add roughly five to ten points; heavy background noise adds five to fifteen.

Then there is the second accuracy problem nobody advertises: speaker attribution. Splitting a recording into "who spoke when" is called diarization, and it is measured separately as diarization error rate. Production systems target under 10%, but on spontaneous conversation with people talking over each other, published results routinely land in the 10–20% band. In practice this means the transcript text may be nearly right while the speaker labels are wrong — which is exactly the failure mode that turns a transcript into a liability if you quote it without checking.

Practical reading of these numbers: at 10% WER, a one-hour meeting produces roughly 8,000–9,000 words, of which around 800–900 are wrong. That is a usable draft and a searchable archive. It is not a document you sign, forward to a client, or paste into a contract without reading.

Language Support Is Not a Yes or No Question

Every vendor's language page is a list of checkmarks. The list is true and almost useless, because accuracy inside that list varies by a factor of five.

Speech models are trained on whatever audio exists in a language, and that is wildly unequal. English, Spanish, German and French have enormous training corpora. Turkish, Vietnamese, Ukrainian and most of Africa and South Asia have far less. On the same model:

  • German typically lands in the 5–7% WER band, close to production quality with light editing.
  • Turkish measures higher and needs review. A 2024 study in the journal Electronics evaluating Whisper for Turkish ASR recorded word error rates of 4.3% to 14.2% across model sizes before fine-tuning, and reported reductions of up to 52% after adapting the model with LoRA.
  • Low-resource languages can exceed 25% WER, which is closer to "notes about the audio" than a transcript.
Turkish is a good illustration of why. It is agglutinative: a single root takes stacked suffixes, generating far more distinct word forms than English does. That drives up out-of-vocabulary words and, with them, error rate. Any language with rich morphology or limited training data behaves the same way.

Two rules follow. First, always test a tool on your own audio in your own language before buying — a ten-minute sample answers more than any comparison table, including this one. Second, if your meetings are genuinely multilingual, transcription and translation should be one step rather than two, because feeding a flawed transcript into a translation engine multiplies both error rates.

File Formats, Length Limits and What Breaks

Most failed transcription jobs fail before the model starts. The file itself is the problem.

FormatWhat it usually comes fromApproximate size for 1 hourNotes
WAVRecorders, professional capture300–600 MBUncompressed; best quality, worst for upload limits
MP3Phone apps, voice memos, podcasts25–60 MBThe safe default; 128 kbps is plenty for speech
M4AiPhone Voice Memos, many Android recorders20–50 MBWidely supported; convert if a tool rejects it
MP4 / MOVScreen and call recordings500 MB–2 GBVideo; extract the audio first if you hit a size cap
OGG / OPUSMessaging apps, browser recordings15–40 MBVery efficient, occasionally unsupported

Three limits catch people out. Monthly minute caps on bundled tools are real: office-suite transcription is typically capped in the low hundreds of uploaded minutes per user per month, which is two or three long meetings. Per-file size or duration caps are common on free tiers and usually sit between 25 MB and a couple of hours. And video files waste both, since you are uploading a hundred times more data than the speech requires.

If a file is rejected, the fix is almost always the same: export audio only, as MP3 at 128 kbps mono. Speech does not need stereo, and a mono file halves your upload.

How to Transcribe a Recording Step by Step

The workflow below applies to any of the five approaches. It takes about fifteen minutes for a one-hour recording, most of which is unattended.

  1. Confirm you have the right to transcribe the file. Consent rules are covered in the next section and are not optional.
  2. Convert the file to a supported format. MP3, mono, 128 kbps. Trim silence at the start and end, plus any pre-meeting small talk you do not want in the record.
  3. Set the correct language before uploading. This is the single highest-impact setting. A model told "English" and fed German audio produces confident nonsense, and a model set to auto-detect can switch languages mid-file when someone code-switches.
  4. Enable speaker labels if the tool offers them, and enter the number of speakers if you are asked. Telling the system there are exactly four voices measurably improves attribution over letting it guess.
  5. Add a custom vocabulary list. Product names, company names, acronyms and the surnames of everyone on the call. This takes two minutes and fixes the errors you would otherwise correct one at a time, forever.
  6. Run the transcription and wait. A one-hour file typically returns in three to ten minutes on a cloud service; a local model on a laptop CPU can take longer than the recording itself.
  7. Review against the audio at the timestamps that matter. Do not proofread the whole thing. Search for the numbers, dates, names and commitments, then listen to those thirty-second windows.
  8. Export in the format that fits the destination. Plain text for search and summaries, SRT or VTT for subtitles, DOCX when someone will edit and comment.
Step seven is where most people either save or waste their afternoon. A transcript is a search index first and a document second — treat it that way and the 10% error rate stops mattering, because you only verify the sentences you actually rely on.

Fix the Audio Before You Fix the Transcript

No app recovers information that was never recorded. Every hour spent improving capture returns more than an hour of editing.

  • Get the microphone within arm's reach. Distance is the single biggest accuracy variable. A wired headset beats a laptop microphone across the room by a wide margin, and it beats an expensive microphone used badly.
  • Kill the loop, not the noise. Speakerphone audio in a hard-surfaced room feeds the microphone a blurred copy of itself. Headphones on each participant eliminate the problem entirely.
  • Record each participant separately when the stakes are high. Separate tracks make diarization nearly trivial, because attribution comes from the file structure instead of the model's guesswork.
  • Ask people to avoid talking over each other. Overlapping speech is the dominant cause of both word errors and mislabeled speakers, and it is the one variable your model cannot fix afterward.
  • Say names at the start. Thirty seconds of "Ayşe speaking, Michael speaking" gives you an anchor for correcting speaker labels later.
For a fuller checklist covering cameras, lighting and internet speed alongside microphones, see our guide on testing your camera and microphone before a call.

Uploading a meeting recording to a transcription service is the processing of other people's personal data on a third party's servers. Two separate questions apply, and answering one does not answer the other.

Were you allowed to record? In the United States, federal law and most states follow one-party consent, meaning a participant may record their own conversation. As of 2026, roughly eleven states — including California, Florida, Illinois, Maryland, Massachusetts, Pennsylvania and Washington — are generally treated as all-party consent jurisdictions, where everyone on the call must agree. Several others are unsettled. Cross-border calls make this messier, not simpler, because the strictest applicable rule tends to govern. Verify the current position for the jurisdictions involved rather than relying on a list, this one included.

In the EU, the GDPR requires a lawful basis under Article 6 before you record, plus transparency about purpose, retention and access — which in practice means telling people in the invitation, not at minute forty. Germany adds a criminal-law layer: § 201 StGB protects the confidentiality of the spoken word, so recording a private conversation without permission is not merely a compliance problem.

Where does the audio go, and for how long? The questions worth asking any vendor:

  • Is the audio used to train the provider's models? A business plan should say no in writing.
  • Is there a data processing agreement, and where are the servers located?
  • Is the original audio deleted after transcription, or retained indefinitely by default?
  • Does the service build voice profiles? Voice used for identification is biometric data and carries stricter obligations.
If the answers are unsatisfying and the content is sensitive, run an open-source model locally. It is slower and less convenient, and the file never leaves your machine. That trade is often the right one for legal, HR and medical recordings — the same logic behind GDPR-compliant video conferencing choices.

What Transcription Costs in 2026

Pricing splits cleanly along the automated-versus-human line, and the gap is roughly twentyfold.

OptionTypical 2026 priceCost of one 60-minute meetingWhat you actually get
Bundled office transcriptionIncluded, minutes capped$0 until the capRaw text, weak speaker separation
Cloud AI, pay per minute$0.05–$0.25 per audio minute$3–$15Text, timestamps, speaker labels, exports
Cloud AI, monthly plan$10–$30 per user per monthEffectively $0 at volumeSame, plus search and storage
Local open-source modelFree software$0 plus your hardware timeFull control, no upload
Human transcription$1–$3 per audio minute$60–$18099%+ accuracy, verbatim options
Platform with built-in transcriptIncluded in the plan$0Transcript exists without a second workflow

The market context: Grand View Research projects the speech-to-text API market will reach roughly $8.6 billion by 2030, growing at about 14% annually. Competition at that scale keeps per-minute pricing falling, which is precisely why paying human rates for routine internal meetings is hard to justify — and why paying human rates for a deposition still is.

Run the arithmetic on your own volume before choosing a plan. Four hours of recordings a month costs $12–$60 on pay-per-minute and $10–$30 on a subscription; twenty hours a month makes the subscription obvious. And if the transcript comes from the meeting platform, the honest comparison is against $0, which is worth checking on your current plan before you add another subscription.

Which Tool Fits What You Record

Four situations cover almost everyone.

Occasional single-speaker files. Voice memos, dictated notes, a recorded lecture. Use whatever is bundled with the software you already pay for and stop shopping.

Regular multi-speaker meetings in one language. A cloud AI service with speaker labels and custom vocabulary earns its subscription within a month, mainly through search rather than reading.

Confidential recordings. Run a model locally, or use a platform that transcribes at the source and deletes the audio. The convenience you give up is the point.

Multilingual meetings. This is where post-hoc transcription is the wrong shape of solution. Transcribing an hour of mixed Turkish, German and English into one text file and then translating it compounds errors at both stages. Capturing each speaker's words in every participant's language while the meeting happens produces a better transcript and removes the file-handling step — which is what Meeyra's meeting translation is built to do across 42+ languages, straight from the browser.

Whichever row you land on, test it on your own worst recording rather than your best one. The tool that handles four people on laptop microphones is the tool that will still be installed in six months. Start a free meeting and check the transcript it produces before you pay for one somewhere else.

Frequently Asked Questions

What is the most accurate speech-to-text app for meetings?

Human transcription remains the most accurate option at 99%+ on decent audio, at roughly $1–$3 per audio minute. Among automated tools, top cloud AI services and locally run open-source models perform within a few points of each other, at about 8–12% word error rate on real multi-speaker meeting recordings.

Can I convert an audio recording to text for free?

Yes. Office suites include transcription with a capped number of uploaded minutes per month, several web tools offer a free tier of 30 to 300 minutes, and open-source models run on your own computer at no software cost. Free tiers usually limit file length, speaker labels and export formats rather than raw accuracy.

How long does it take to transcribe one hour of audio?

A cloud AI service typically returns a one-hour file in three to ten minutes, since it processes segments in parallel. A local model on a laptop without a GPU can take as long as the recording itself. Manual transcription by a person runs at a ratio of about four to six hours of work per hour of audio.

Which audio format works best for transcription?

MP3 at 128 kbps in mono is the best default: universally supported, small enough for upload limits, and lossless enough for speech. WAV offers marginally better quality at ten times the file size, and video files should have their audio extracted first to avoid size caps.

Does speech-to-text work well in languages other than English?

It works well for high-resource languages such as German, Spanish and French, which land near 5–7% word error rate. Accuracy drops for languages with less training data or complex morphology — published research measures Turkish between 4.3% and 14.2% depending on model size — so always test with your own audio before committing.

Do I need everyone's permission to record and transcribe a meeting?

In most US states one participant's consent is enough, but around eleven states require consent from everyone on the call, and the EU requires a lawful basis plus advance transparency about purpose and retention. The safe default in every jurisdiction is to announce the recording in the invitation and again at the start of the meeting.

Can transcription tools tell speakers apart?

Most cloud services attempt it through diarization, and accuracy is meaningfully lower than word accuracy — error rates of 10–20% are common in conversations with crosstalk. Telling the tool how many speakers are present, or recording each participant on a separate track, improves results substantially.

Should I transcribe the recording or use live transcription instead?

If you control the meeting platform, live transcription is usually the better path: the transcript exists the moment the call ends, no file is uploaded anywhere, and speaker attribution comes from the platform rather than from a model guessing. Post-hoc transcription is for audio you did not capture yourself, such as phone recordings, interviews and in-person sessions.