Personal utility · macOS · Open source · formerly Voice Terminal
Overtone: Dictate Straight Into Claude Code
A small macOS menubar utility for talking to your terminals. Hold the right Option key, say what you want, let go, and the words paste into whichever window has focus. I built it because typing had become the slowest part of running several Claude Code sessions at once, and it has turned into the personal tool I reach for most. This is what it does, how it is put together, the details that made it reliable enough to forget about, and why small utilities like this belong in every builder's workflow.
TL;DR
- What it is. A few hundred lines of Python living in the macOS menubar. Hold Right Option, speak, release, and the transcript pastes into the focused app. Hold Right Command to send your clipboard plus a spoken question to Claude and paste the answer. Works in any app, built for terminals.
- Why it matters. Speech runs roughly three times faster than typing, and the drop in mental fatigue is bigger than the speed gain. When you are steering three or four agent sessions, your job is intent, and intent is far easier to say than to type.
- How it is built. rumps for the menubar, pynput for the global hotkey, sounddevice and numpy for audio, Whisper for transcription, NSPasteboard through PyObjC plus a synthesized ⌘V for the paste. One file, no build step.
- What made it disappear. Modifier-only hotkeys, a hold delay, shortcut cancellation, a full clipboard snapshot and restore, silence gating, main-thread UI, and a busy guard. Each one removed a specific daily annoyance.
- The pattern. A global hotkey, a slice of desktop state, one API call, and a paste. That shape covers an enormous range of personal tools, and every builder should have a few of them.
01Why talk to a terminal
Most of my hands-on time now has the same shape. Three or four Claude Code sessions are open across different repos, each working on something that takes minutes to finish. My job in that setup is direction. Each message is a short paragraph of intent, what to do, what to check, what to leave alone. The agent does the typing that used to be the work. What remains for me is composing those paragraphs, and it turns out composing paragraphs on a keyboard, four terminals at a time, is where the hours and the fatigue go.
Speaking fixes both. Conversational speech runs around 150 words a minute against 40 to 70 for most people typing, so the raw throughput roughly triples. The larger effect is the one that is harder to measure. Saying a sentence out loud costs almost nothing, so the extra clause gets included. "Run the suite, and if the timeout test fails, check whether the fixture still points at the old port" is a sentence I will say every time and type maybe a third of the time. Agents do better with more context, and voice makes context cheap. At the end of a long day of orchestration, the difference in how tired I am is not subtle.
The obvious answer is the dictation built into macOS, and it is a poor fit here. It behaves unreliably in terminal apps, it has no push-to-talk, and it wants to own the whole input rather than sit alongside the keyboard. What I wanted was one key that I could hold like a walkie-talkie, in any window, with the result arriving as a single paste. That is the entire product.
02Hold, speak, release
The app sits in the menubar as a small microphone. Focus the window you want to talk to, hold the hotkey, wait for a short tink, say your piece, and let go. A pop confirms the recording ended, the icon flips to an hourglass while Whisper works, and a second later the transcript pastes at the cursor. Your clipboard is put back the way it was. Two modes share the same motion.
| Hold | What happens | Menubar |
|---|---|---|
| ⌥ Right Option | Record while held. On release, transcribe and paste the words into the focused window. | 🔴 recording ⏳ transcribing |
| ⌘ Right Command | Snapshot the clipboard, record while held. On release, send the clipboard as context plus the spoken request to Claude, and paste the response. | 🟣 recording 🤖 waiting on Claude |
The multi-session flow in the animation above is the reason the tool exists. Click the first pane, hold, "run the eval suite on the new minimal pairs and tell me which ones regressed", release. Click the second pane, hold, describe the next thing, release. By the time the third instruction is in, the first session is already reading files. Nothing in that loop requires a keyboard beyond the one key under your thumb, and the terse-command habit that a keyboard encourages goes away with it.
03What it is built from
Everything lives in one Python file, a little over 450 lines, with no build step and nothing installed system-wide. The libraries each do one job.
- rumps wraps the Cocoa status bar so the whole app is a class with a title, a menu, and a run loop. The icon is just the title string, which is why the states are emoji.
- pynput provides the global keyboard listener. It reports Right Option and Right Command as distinct keys on macOS, which the hotkey design depends on.
- sounddevice opens a 16 kHz mono input stream and hands audio to a callback in float32 chunks. numpy concatenates the chunks, measures loudness, clips, and converts to 16-bit samples for a WAV file built in memory with the standard library.
- openai is the client for Whisper. The base URL and model are configurable, so the same code talks to OpenAI directly or to any OpenAI-compatible endpoint, including proxies that front faster Whisper variants. The same client shape handles the Claude call in the second mode.
- PyObjC gives direct access to
NSPasteboardfor reading, snapshotting, and restoring the clipboard, and toAppHelper.callAfterfor getting work back onto the main thread. - osascript sends the ⌘V. One line of AppleScript through System Events, which is also why the app needs Accessibility permission.
- python-dotenv reads the keys from a local
.env. That file is the entire configuration surface.
Whisper prices at about six tenths of a cent per minute of audio, so a five-second instruction costs a small fraction of a cent, and a heavy day of dictation runs to pennies. Through a proxy that serves a turbo variant, a short clip comes back in well under half a second, which is close enough to instant that the paste feels like it followed the release of the key.
04The details that made it disappear
The first version worked in an evening and was annoying for months. Every fix below removed one specific daily irritation, and together they are the difference between a demo and a tool you stop noticing.
- Modifier-only hotkeys. The original hotkey was ⌘⇧Z, which is Redo in nearly every app, so each dictation also undid an undo somewhere. ⌥Space typed a space into the terminal that could not be suppressed. A modifier held on its own sends nothing to the focused app, so there is nothing to suppress. Using the right-hand keys leaves the left ones free for normal shortcuts.
- Hold delay and shortcut cancellation. Recording starts 150 ms after the key goes down, so a tap does nothing. If any other key is pressed while the hotkey is held, the gesture is treated as a shortcut, ⌥← or ⌘C, and the recording is cancelled or never started. Clips under 0.3 s are dropped as accidental.
- The twin-modifier quirk. pynput derives modifier presses and releases from a flag mask, so releasing Right Option while Left Option is still down is reported as a second press of Right Option. Modifiers never auto-repeat, so a repeated press of the held key is treated as its release. Without this the recording would run until the next keystroke.
- Ignore your own keystrokes. The ⌘V the app sends shows up in its own listener. pynput flags injected events, and the callbacks return early on them.
- Clipboard as the transport. Typing the transcript character by character was slow and locked focus to the target window for the duration. Pasting is instant. The app snapshots every item and data type on the pasteboard, writes the transcript marked with the
org.nspasteboard.TransientTypeflag so clipboard managers skip it, sends ⌘V, waits half a second for the target app to read asynchronously, and restores the snapshot only if the change count has not moved. Copy something mid-dictation and it wins. - Silence gating. Whisper hallucinates on silence, most often the phrase "Thank you." The app measures RMS over 50 ms windows and drops a recording whose loudest window sits below a threshold near -46 dBFS. The threshold is deliberately conservative, since dropping real speech is worse than an occasional stray phrase.
- Threads and the main thread. The key listener, the hold timer, and the transcription worker each run on their own thread and share a small state machine, idle to recording to processing, behind one lock. Menubar title, status, and notifications are AppKit and only touched from the main thread through
callAfter. Before that discipline the app would crash a few times a week in ways that were hard to reproduce. - Callbacks that never raise. An exception escaping a pynput callback stops the listener for good, silently. Every callback catches everything and turns it into a notification.
- Busy guard. Holding the hotkey while the previous clip is still processing plays a third sound and does nothing else, so two dictations never race for the same paste.
None of these are clever. They are the kind of thing you only find by using a tool daily and refusing to live with the paper cuts, and they are exactly the work that a general-purpose product cannot do for your particular hands.
05A model one key away
The second hotkey grew out of a habit. I would copy an error out of a terminal, switch to a chat window, paste it, type a question, copy the answer, switch back. Right Command collapses that loop into the same gesture as dictation. Copy something, hold the key, ask, release. The clipboard is captured at the moment the key goes down, the spoken request is transcribed, and both go to Claude with a system prompt that says the reply will be pasted into an editor or terminal, so keep it direct. The response lands where the cursor is.
Some of what that gets used for in a normal day:
- Copy a stack trace, hold, "what is wrong here and what would you check first". The answer arrives in the terminal beside the trace.
- Copy a diff, hold, "write the commit message". Paste lands in the commit editor.
- Copy a function, hold, "add type hints and a docstring", with the cursor on the original.
- Copy a long thread, hold, "draft a reply that agrees with the second option and asks about the timeline".
The value comes from asking the question at the place the answer is needed, with the relevant context already attached, while no window changes.
06Build your own
Strip the specifics away and the shape of this tool is generic. A trigger you can reach without looking, one slice of desktop state, one API call, and one action that puts the result back where you are. Overtone binds a held key to the microphone, the clipboard, a transcription model, and a paste. Almost every part of that sentence can be swapped.
- Other slices of state. The current selection, the focused window's title, a screenshot of the frontmost window, the URL in the browser, the file open in the editor. Each one is a few lines on macOS and each one is context a model can use.
- Other actions. Paste is the simplest. Pressing Return after the paste turns dictation into dispatch. A notification, a new note, a clipboard write, or a command run in a specific terminal are all one call away.
- Where this one goes next. A job queue that remembers which window each recording was aimed at, so you can dictate into three terminals back to back without waiting for the paste. A local Whisper for airplanes. A spoken "and submit" that ends the instruction and presses Return.
The tools worth having are the ones shaped exactly to how you work, and the cost of building them has collapsed. A utility like this takes a weekend, and the reliability pass takes an afternoon here and there for as long as you keep using it. Everyone who spends their day driving agents should have a few of these, small programs that sit between their intentions and their machine and remove a little friction each. This one removed the keyboard.
Overtone is open source at github.com/nbdesai1992/overtone. It needs macOS, Python 3.9 or newer, an API key for Whisper, and two permissions on first run, Microphone and Accessibility. Setup is a clone, a virtualenv, and one line in a .env file.
# clone, install, configure, run
git clone https://github.com/nbdesai1992/overtone.git && cd overtone
python3 -m venv venv && source venv/bin/activate && pip install -r requirements.txt
echo "OPENAI_API_KEY=sk-your-key" > .env
python overtone.py # 🎤 in the menubar. Hold Right Option and talk.