Overview
Speech Notes is a browser-based AI speech to text workspace that turns recordings, media files and live conversations into editable, searchable transcripts. Speech Notes keeps audio intake, speech recognition, transcript review and export inside one workspace, so spoken material reaches a finished deliverable without moving between an uploader, an editor, a search tool and a converter.
A transcript is only useful when it is easy to check and reuse. Speech Notes therefore treats recognition output as a working draft rather than a finished record: names, figures, overlapping voices and specialist vocabulary are expected to need correction, and the editor, search, speaker controls and export options sit after recognition for exactly that reason.
Key Features
- Three intake paths: select an audio or video file already on the device, capture a fresh recording in the browser, or import a supported public media URL.
- Broad format support: MP3, WAV, M4A, AAC, WEBM, MP4 and MOV, with media intake up to 1GB per file.
- Transcription across 100+ languages powered by OpenAI Whisper, with automatic language detection or explicit language selection.
- A wider model lineup including Nemotron 3.5 ASR, Seed Audio 1.0, Miso One, OmniVoice and VoxCPM.
- Speaker labels that clarify multi-person recordings, with generic speakers renameable during the editorial pass.
- A transcript editor with language-aware search: locate a person, term or decision, replay only the nearby audio, and confirm the wording before it is quoted.
- Purpose-specific exports: TXT for plain text, SRT and VTT for timed captions, DOCX and PDF for readable documents, and JSON for structured data workflows.
- AI Summary, AI Analytics and Chat with AI for interrogating recognized content, plus transcript translation across 100+ languages and email delivery on paid plans.
Use Cases
Meetings and calls: preserve the discussion, then extract owners, decisions, objections and unresolved questions into a record the team can search. Interviews and research: retain speaker context and verify any quotation that will appear in published work. Podcasts and creator media: work from the final cut to prepare searchable copy, caption files and editorial source material. Lectures and lessons: convert a spoken explanation into reviewable notes while keeping terminology and worked examples intact. Video captions: use timed output as a draft, then check reading speed, line breaks, names and synchronization before publishing. Operational records: make customer calls, project updates and discovery sessions searchable after the event.
Getting Started
Open speechnotes.org in a browser, add a file, record new audio or paste a supported media link, confirm the language and speaker options, and start the transcription. Review the result in the editor, then export the format the next task expects.
Pricing and Plans
Speech Notes is freemium. Every account starts with 5 trial minutes alongside free browser-native Whisper transcription. Starter costs $4.90 per month billed as one annual charge of $58.80 and includes 1,440 transcription minutes per year. Pro costs $14.90 per month billed annually at $178.80 for 7,200 minutes per year and adds AI Summary, AI Analytics, Chat with AI and translation. Max costs $24.90 per month billed annually at $298.80 for 36,000 minutes per year.






