LLM prompt engineering and evaluation

I have spent three years on production LLM work: designing and optimizing prompts at Remotasks, then leading the team at Scale AI that produced and audited the software training data models learn from. The part I care about most is the part people skip. A prompt is something to measure, not something to word well, so I build the regression gate first and change the prompt second.

Hiring for prompt engineering or evaluation? Email me.

Where I've done this work

A prompt I built and measured

I use an LLM to draft and edit all the published copy for One Group Central, a member portal product I build. The copy was factually correct and read as machine written, which on a product whose pitch is trustworthy elections costs the sale.

My first attempt was wrong, and the way it was wrong is the useful part. I assumed a vocabulary problem, gave the model a banned word list, and watched it plateau. The tell is not the words. It is the habits, and habits survive a synonym swap: dashes as the default connector, lists shaped as bold label then colon then explanation, groups of three by reflex, a closing recap. The second thing I got wrong was worse. "Write more naturally" is unfalsifiable, so neither the model nor I could tell whether it had complied, and the behavior drifted back within a session or two.

So I rewrote the instruction as something checkable, moved it out of the chat and into the project's system prompt where every session and every subagent loads it, and then built the thing that makes it stick: a 239‑line checker that extracts prose from HTML, markdown, JavaScript, and Python, applies the rules per line, and exits nonzero on a violation.

What the number does not prove: the checker measures compliance with my rules, not quality. Zero errors means the known tells are gone, not that the writing is good, which is why the gate is mechanical plus a human reread. The validation I have not run is a blind test of before‑and‑after pairs against real readers, and that is the next thing I would build.

How I think about evaluation

Guardrails

The rules I actually care about are enforced outside the model, because a rule a model can be argued out of is a preference and a rule in the deploy pipeline is a constraint. Three that are in production:

What I haven't done

Worth stating plainly, because an engineer who will not tell you the edge of their experience is expensive to hire. I have not built a RAG pipeline or run a vector database in production. I have not implemented an LLM‑as‑judge, though I know how I would validate one. No fine tuning, no RLHF. No production multi‑agent framework, though I orchestrate subagents daily and have hit the real problem that comes with it, which is context isolation between them.

Let's talk

I'm in Mount Sinai, New York, and I work remotely. Available after 12pm Eastern on weekdays.