Neal Desai Staff PM at Scale AI · ex-CTO · evals & agent tooling GH LI

Personal utility · macOS · Open source · formerly Voice Terminal

Overtone: Dictate Straight Into Claude Code

A small macOS menubar utility for talking to your terminals. Hold the right Option key, say what you want, let go, and the words paste into whichever window has focus. I built it because typing had become the slowest part of running several Claude Code sessions at once, and it has turned into the personal tool I reach for most. This is what it does, how it is put together, the details that made it reliable enough to forget about, and why small utilities like this belong in every builder's workflow.

overtone on GitHub How it is built Build your own

TL;DR

Three sessions, one key. Click a pane, hold Right Option, speak, release. The transcript lands in one paste and the session gets to work while you move to the next.

01Why talk to a terminal

Most of my hands-on time now has the same shape. Three or four Claude Code sessions are open across different repos, each working on something that takes minutes to finish. My job in that setup is direction. Each message is a short paragraph of intent, what to do, what to check, what to leave alone. The agent does the typing that used to be the work. What remains for me is composing those paragraphs, and it turns out composing paragraphs on a keyboard, four terminals at a time, is where the hours and the fatigue go.

Speaking fixes both. Conversational speech runs around 150 words a minute against 40 to 70 for most people typing, so the raw throughput roughly triples. The larger effect is the one that is harder to measure. Saying a sentence out loud costs almost nothing, so the extra clause gets included. "Run the suite, and if the timeout test fails, check whether the fixture still points at the old port" is a sentence I will say every time and type maybe a third of the time. Agents do better with more context, and voice makes context cheap. At the end of a long day of orchestration, the difference in how tired I am is not subtle.

The obvious answer is the dictation built into macOS, and it is a poor fit here. It behaves unreliably in terminal apps, it has no push-to-talk, and it wants to own the whole input rather than sit alongside the keyboard. What I wanted was one key that I could hold like a walkie-talkie, in any window, with the result arriving as a single paste. That is the entire product.

02Hold, speak, release

The app sits in the menubar as a small microphone. Focus the window you want to talk to, hold the hotkey, wait for a short tink, say your piece, and let go. A pop confirms the recording ended, the icon flips to an hourglass while Whisper works, and a second later the transcript pastes at the cursor. Your clipboard is put back the way it was. Two modes share the same motion.

HoldWhat happensMenubar
⌥ Right Option Record while held. On release, transcribe and paste the words into the focused window. 🔴 recording ⏳ transcribing
⌘ Right Command Snapshot the clipboard, record while held. On release, send the clipboard as context plus the spoken request to Claude, and paste the response. 🟣 recording 🤖 waiting on Claude

The multi-session flow in the animation above is the reason the tool exists. Click the first pane, hold, "run the eval suite on the new minimal pairs and tell me which ones regressed", release. Click the second pane, hold, describe the next thing, release. By the time the third instruction is in, the first session is already reading files. Nothing in that loop requires a keyboard beyond the one key under your thumb, and the terse-command habit that a keyboard encourages goes away with it.

03What it is built from

Everything lives in one Python file, a little over 450 lines, with no build step and nothing installed system-wide. The libraries each do one job.

Hold Right ⌥ pynput · 150 ms hold Record sounddevice · 16 kHz Release → WAV numpy · clip · gate Whisper OpenAI-compatible URL Paste NSPasteboard · ⌘V Hold Right ⌘ same listener Read clipboard captured on press Claude context + request Claude mode joins the same recording path, then detours through the model before the paste
One recording path. The two hotkeys differ only in what is captured before and what happens after Whisper.

Whisper prices at about six tenths of a cent per minute of audio, so a five-second instruction costs a small fraction of a cent, and a heavy day of dictation runs to pennies. Through a proxy that serves a turbo variant, a short clip comes back in well under half a second, which is close enough to instant that the paste feels like it followed the release of the key.

04The details that made it disappear

The first version worked in an evening and was annoying for months. Every fix below removed one specific daily irritation, and together they are the difference between a demo and a tool you stop noticing.

None of these are clever. They are the kind of thing you only find by using a tool daily and refusing to live with the paper cuts, and they are exactly the work that a general-purpose product cannot do for your particular hands.

05A model one key away

The second hotkey grew out of a habit. I would copy an error out of a terminal, switch to a chat window, paste it, type a question, copy the answer, switch back. Right Command collapses that loop into the same gesture as dictation. Copy something, hold the key, ask, release. The clipboard is captured at the moment the key goes down, the spoken request is transcribed, and both go to Claude with a system prompt that says the reply will be pasted into an editor or terminal, so keep it direct. The response lands where the cursor is.

Some of what that gets used for in a normal day:

The value comes from asking the question at the place the answer is needed, with the relevant context already attached, while no window changes.

06Build your own

Strip the specifics away and the shape of this tool is generic. A trigger you can reach without looking, one slice of desktop state, one API call, and one action that puts the result back where you are. Overtone binds a held key to the microphone, the clipboard, a transcription model, and a paste. Almost every part of that sentence can be swapped.

The tools worth having are the ones shaped exactly to how you work, and the cost of building them has collapsed. A utility like this takes a weekend, and the reliability pass takes an afternoon here and there for as long as you keep using it. Everyone who spends their day driving agents should have a few of these, small programs that sit between their intentions and their machine and remove a little friction each. This one removed the keyboard.


Overtone is open source at github.com/nbdesai1992/overtone. It needs macOS, Python 3.9 or newer, an API key for Whisper, and two permissions on first run, Microphone and Accessibility. Setup is a clone, a virtualenv, and one line in a .env file.

# clone, install, configure, run
git clone https://github.com/nbdesai1992/overtone.git && cd overtone
python3 -m venv venv && source venv/bin/activate && pip install -r requirements.txt
echo "OPENAI_API_KEY=sk-your-key" > .env
python overtone.py          # 🎤 in the menubar. Hold Right Option and talk.