Neal Desai

Staff PM at Scale AI · ex-CTO · evals & agent tooling

Staff Strategic PM at Scale AI, working with frontier labs on RL environments, fine-tuning data, and model evals. Before that: startup co-founder and CTO, production ML at Beyond Limits, predictive models at Epic. Caltech and UW-Madison. New York.

Benchmarks

Measuring where agents fall short, with a live leaderboard.

Methodology

Playbooks for building evals and agent skills that hold up.

02 Synthetic Artifacts and Evaluation EnvironmentsWriteup Test agents on work that looks real. Synthetic artifacts that pass for real, synthetic MCP servers that behave like the real ones, and tests built from known answers, all made by a quality loop that never grades its own work. Read the writeup →
03 Loop Engineering for Agent SkillsWriteup Skill triggering, measured as a compliance rate. Synthesize positives and hard negatives per skill, run them through Claude Code or Codex, score invocations straight from the tool log, and hill-climb with targeted edits until the curve flattens. Read the writeup →

Tools

Open-source tools I use every day alongside Claude Code, from a single keystroke to a whole app.

04 OvertoneOpen source · macOS Hold a key, speak, release. The words land in the focused terminal. A few hundred lines of Python in the menubar, plus a second key that sends your clipboard and a spoken question to Claude and pastes the answer. The tool I reach for most. Read the writeup →
05 RelayOpen source · v0.5 · in progress Label a Todoist task @claude. Claude Code does it on your machine. A Claude Code plugin: one worker per task, in parallel, with results posted back as review comments. A guard keeps workers read-only against the outside world. No server, no API key, your own Claude plan. Read the writeup →
06 ConsulOpen source · Render Describe an app in plain words. Consul plans it, builds it, and puts it online. A Claude Code toolkit that turns one sentence into a brief you approve, then hands the work to three specialists and a separate checker until the app is live on Render and every requirement has evidence. Everything runs on your own GitHub and Render accounts. Read the writeup →

I'm a Staff Strategic Product Manager at Scale AI, a pre-sales role working with research teams at frontier AI labs on reinforcement-learning environments, fine-tuning data strategy, and model evaluations. Much of the work is scoping and building data products and the unique insights that come with them, from studying where models fall short and closing the gap.

Before Scale I was a startup co-founder and CTO, built production ML at Beyond Limits for enterprise customers in energy, finance, and healthcare, and started out at Epic Systems shipping predictive models for hospitals. I studied biological engineering and applied math at Caltech and machine learning at UW-Madison. I'm based in New York. Outside of work I mentor high school students through their first real research projects in healthcare and tech, which is some of the most rewarding work I do. The rest of the time you'll usually find me traveling or outdoors.

Have an idea, a question, or just want to say hi? I'd love to hear from you. Reach me on LinkedIn or by email.