AI News HubLIVE
サイト内リライト7 分で読了

翻訳待ち:My 3 favorite AI tools for voice dictation while vibe coding - and one is free

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:I've tested the top AI voice tools to find which delivers the best accuracy, privacy, and corrections.

ソースZDNet AI

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。

Follow ZDNET: Add us as a preferred source on Google. ZDNET's key takeaways Fast self-correction matters more than raw accuracy alone.Wispr Flow led on corrections, vocabulary, and reliability.Free, local FluidVoice nearly matched the paid winner.Since February, I have dictated 120,896 words. In the last three months alone, I have dictated 50,234 words. I've done this across 2,314 individual microphone recording sessions, averaging about 24 words per dictation sequence. With improvements in AI, dictation quality, speed, and accuracy have come a long way. Over the years, I've tried to work with speech recognition many times, but it hasn't been very successful until quite recently. A typical article is roughly a thousand words. So if you look at it that way, I have dictated the equivalent of roughly 120 articles since February. Also: 7 surprisingly useful ways to use ChatGPT's voice mode, from a former skepticThat's not to say I don't type. I type a lot, but it's clear that I also use dictation a lot. I've adopted voice dictation as a primary input modality for two key reasons. First, it helps to protect my wrist, which tends to have carpal tunnel symptoms. By dictating, I'm using my wrist a little bit less. In this article, I'll show you the three primary contenders that I've looked at and spent time with this year. I'm also going to go over a number of also-rans and honorable mentions, just to give you an idea of the dictation tools that are out there and how they differ. Where voice dictation fits into my workflow I use voice dictation a lot when working with Claude Code or OpenAI's Codex to vibe code any of the products I'm working on. I use it a lot in Slack and Google Chat when talking to my editors and some of my project partners. Also: I built an iOS app in just two days with just my voice - and it was electrifying I use voice dictation quite a lot when replying to email messages. I use it a little bit less when composing new messages, but I use it there sometimes as well. I use it a lot when taking notes. I also use it sometimes for search phrases in Google or prompts that I give to ChatGPT, Gemini, or Claude Code. (Disclosure: Ziff Davis, ZDNET's parent company, filed an April 2025 lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.)Also: I built two apps with just my voice and a mouse - are IDEs already obsolete? I find that voice dictation is sometimes a little bit less precise and a little bit less disciplined than writing each word with the keyboard. But I can dictate at an average of 118 words per minute, while my typing is usually in the 70- to 80-word-per-minute range. I am dictating most of this article just as a proof of concept to make the point that an article can be dictated. But that's not the norm for me. Despite all the dictation I do, I don't usually dictate my articles. I sometimes dictate a few lines of an article, but mostly I type them out. I sweat each sentence, choosing the words and structure with great care. Typing lends itself to that kind of attention to detail.But when I'm vibe coding, for example, and I'm discussing how I want a feature to be instantiated, how I want something to behave, or a bug that I've noticed, speaking is considerably faster and also considerably more gentle on my hands than typing it into the computer. The features that matter most I find two features to be mission-critical. The first is a customizable dictionary, so that when I say something like ZDNET, the dictation product understands how I want it spelled and presented. The second is an on-the-fly correction capability, so that when I say something and then correct myself and re-say it, the version that lands in whatever I'm dictating into contains those corrections.Also: I'm an AI tools expert, and these are the 4 I pay for now (plus 2 I'm eyeing)Almost all the dictation products are initiated by a hotkey. I bind the dictation hotkey to a button on my mouse so that when I tap the button, the dictation starts or stops. This allows me to dictate regardless of what application or web page I'm in at the moment. It means I can do computer input even if my keyboard isn't in front of me. To that end, I will be spotlighting three products: , , and 1. Wispr Flow: Best overall At $144 a year, or $15 a month, is certainly not cheap, but I would argue it's actually worth it. Despite trying almost all the other products, this is the one I keep coming back to and have used more than any other. Wispr Flow is the only one of our top three available for Mac, Windows, iOS, and Android. It is not, however, available on Linux, although the company has a waitlist for Linux users. Wispr Flow's standout feature, at least in terms of my usage, is its in-flight self-correction. As you're dictating, you can correct yourself, and it updates what's being transcribed. Once you get used to this feature, you really don't want to go back, especially if you're doing a large amount of dictation like I do. It means that the text you produce is, more often than not, usable because if you misspeak, you can fairly easily correct it as you're speaking and end up with a decent result. None of the other models that I tested were able to do this as smoothly. Some couldn't do it at all. FluidVoice has come close, but I would say that Wispr Flow made accurate corrections eight out of 10 times, and FluidVoice made accurate corrections maybe four out of 10 times. For in-flight self-correction, that's measurable when you're doing a lot of work. Also: I tested 3 text-to-speech AI models to see which is best - hear my resultsI also found that Wispr Flow's dictionary is reliable and effective. What I mean by that is that once I've trained it on an incorrectly spelled word or incorrectly interpreted word, I almost never have to go back and correct it again. Once I trained it on the word ZDNET, for example, Wispr Flow reliably gets it correct just about 100% of the time. That's also the case with my library of 90 or so other words that I regularly correct. Wispr Flow has two dictionary options: It allows you to feed it individual words like Gewirtz, and it allows you to feed it misspellings or misinterpretations and then the corrected word. For example, it regularly had trouble with the word Claude, which it would represent as "call it." I set up a dictionary definition for "call it code" that converted to Claude Code, and I've never had a problem since. Once in a while, Wispr Flow misses the insertion of a chunk of text into the destination location. For example, I might dictate a paragraph that I want to go into Notes, and it never winds up there. Wispr Flow keeps a history of dictation in its app. If it misses insertion, I can open it up in the app, copy from the history, and paste it in. I don't ever actually lose any of my dictation, even if it doesn't always arrive on target the first time out (which is a fairly rare occurrence). Beyond price, my biggest concern about Wispr Flow is that it's a cloud-only model, meaning that all of your voice snippets are sent to the cloud for transcription. Despite the similarity in names, Wispr Flow is not based on OpenAI's open-source Whisper speech recognition technology. Wispr Flow appears to be its own model or based on a stack of a variety of model providers. The company does not disclose the exact model used. Also: I tested ChatGPT's Live Voice upgrade, and it almost felt human - how to try itWispr Flow offers a number of data and privacy options in its settings, including a privacy mode, the option to turn private cloud sync on and off, and local data storage. However, what the options are called in the UI and what the options actually do are different. Privacy mode isn't really what you would think. It's not that it doesn't look at any of your phrases. It's that when turned on, it will not send any of your data to be used for training the AI. Private Cloud Sync, when turned off, does not mean that the data is not sent up to the cloud. It means that it's not stored in the cloud to sync to other devices. It is still sent up to the cloud for transcription, but Wispr Flow then deletes the data immediately after transcription. The local data storage option does not control whether data is stored locally or in the cloud, but instead controls factors like whether or not Wispr Flow will auto-delete local data every 24 hours or never store any data locally, meaning, for example, that the dictation history would not be available to you. Also: I used Gmail's AI tool to do hours of work for me in 10 minutes - with 3 promptsIf you have data control policy concerns, confidentiality concerns, disclosure restrictions, or any other legal reason you don't want your data up in the cloud, you might want to avoid Wispr Flow. I have found, for basic productivity, that Wispr Flow has become my most actively used voice dictation product. I have been cycling through a bunch of them to try to find one that I could live with as a daily driver. So far, that's Wispr Flow, and that's why it's my top recommendation. Wispr provided me with a Pro account to use for a year for evaluation, but there's a very good chance that when that year runs out, I will renew it with my own money. That should tell you something. There is a trial version of Wispr Flow that allows you to use it for up to 2,000 words. 2. Superwhisper: Best for voice dictation customization Our second tool, , is just chock-full of features. Even so, I just really haven't been able to mesh with it. I am presenting it here because it does have so many specialized capabilities that you may find helpful. I found the lack of on-the-fly correction to be a deal killer. To be honest, that surprised me because I didn't even realize that I had been using the on-the-fly correction as actively as I was until it became apparent that when it was missing, I missed it greatly. Superwhisper is available for $8.49 a month, $84.99 a year, or a one-time purchase of $249 that provides unlimited lifetime use. If you think you're going to be using it for a number of years, that's a good deal. On the other hand, AI-based solutions are changing so rapidly that there may be a far better solution, a far cheaper solution, or a free solution available before you fully utilize the unlimited lifetime use. There is a free version that you can use with smaller voice recognition models. Also: Which AI tools are actually worth paying for? I'm keeping these subscriptions in 2026 - here's whyI should point out that recognition accuracy is not really that much of an issue between the two top products. Both are able to run models that are successful for overall recognition. Although Superwhisper provides a lot of model choices, some of them do not perform voice recognition as accurately but have other advantages. Superwhisper's biggest advantage over Wispr Flow is that you can create a Mac-only, offline-only, no-data-in-the-cloud version. If you pay for the one-time lifetime use with no additional billing ever, you can own it and control it all and have all your dictation happen entirely on your computer. Superwhisper will do that for you. Wispr Flow will not. Superwhisper's standout fiddly feature is its mode system. Modes are saved processing pipelines. Basically, what this means is that depending on what mode you're in, Superwhisper can behave or function completely differently. A mode consists of four elements: The voice model, which is a speech-to-text engine that transcribes the spoken word for you.The language model, which is the part that cleans up and reshapes the raw transcript or modifies it in some way.A set of processing instructions, essentially a recipe or a skill attached to that mode.An auto-activation rule, which says that the mode becomes active on a given app or website. For example, you could have a mode that becomes active when you are using Gmail, another [truncated for AI cost control]