After previewing at I/O 2026 in May, Gemini for macOS is now rolling out advanced voice control capabilities.
Google bills this as letting you “speak naturally into any window on your desktop,” with two key features.
The first is intelligent dictation that transcribes “spoken words into clean, polished text” that automatically removes “umm” and “ah,” while accounting for your mid-sentence corrections. The formatted text is entered where your cursor currently is. It should offer a speech-to-text experience that’s similar to Gboard Rambler on upcoming Gemini Intelligence.
The other capability goes beyond transcription by understanding the context of your current screen to perform complex tasks. This is an opt-in capability by enabling Gemini reasoning in settings.
- Extract and summarize information: Highlight local files, images, or documents on your desktop and tell Gemini exactly what you need. For example, you can say “Read these vet files and summarize my dog’s medical history in an email to the kennel.”
- Compose and rewrite text: Highlight text anywhere on your screen and use your voice to instantly rewrite it, tweak the tone, and drop the polished version right where you need it. Try saying, “Turn these notes into an executive summary with a TL;DR at the top.”
- Generate and edit images: Use your voice to create visuals when conceptualizing an idea or enhancing a travel itinerary, or iterate on a design on your desktop by referencing it to request quick updates. Try saying, “Take this illustration and generate a dark-mode version of it.”
Long-press the Fn key wherever you are in macOS to enable, or tap the new screen sharing button at the end of the “Ask Gemini” prompt box.
Make sure you’ve updated to the latest version (1.88) of Gemini for macOS, with this new voice capability rolling out globally to all users in English, with more languages “coming soon.”