Turning speech into accurate, structured text is now a solved-enough problem that the interesting competition happens after transcription: summaries, action items, clinical notes, legal records, and searchable archives. Users span anyone who talks for a living — sales reps and executives drowning in meetings, clinicians documenting visits, lawyers preparing depositions, journalists, and developers who'd rather dictate than type.
Modern tools run speech-recognition models (increasingly fine-tuned per domain) with speaker diarization, then apply LLMs to reshape the raw transcript into whatever the job requires. Granola, which has raised $168M, exemplifies the meeting-native approach — an AI notepad that blends your typed notes with what was said. Otter.ai ($50M) remains a household name for meeting transcription and summaries, while Rev ($99M) serves accuracy-critical legal work by combining AI with human transcriptionists. At the input end, Aqua Voice ($2M) focuses on dictation tuned for technical vocabulary, and vertical players like Andy AI target clinical documentation for home health.
Accuracy on clean audio has commoditized; leaders now differentiate on hard conditions — accents, crosstalk, jargon — and on workflow depth: does the tool push action items into your CRM or EHR, or just email you a summary? Buying considerations include privacy posture (does audio train vendor models?), whether bot-free capture matters for your meeting culture, per-seat versus usage pricing, and domain accuracy — always test with your own recordings. NeuronFeed tracks 15 AI transcription companies with $716M in combined funding.