LLM prompt engineering and evaluation
I have spent three years on production LLM work: designing and optimizing prompts at Remotasks, then leading the team at Scale AI that produced and audited the software training data models learn from. The part I care about most is the part people skip. A prompt is something to measure, not something to word well, so I build the regression gate first and change the prompt second.
Hiring for prompt engineering or evaluation? Email me.
Where I've done this work
- Remotasks (2022 to 2023, Instructor). I designed and optimized production LLM prompts for summarization, classification, and entity extraction, iterating by experiment and keeping the versions that measured better rather than the ones that read better. A large part of the job was adversarial: finding the inputs where a model returns a confident wrong answer, and closing that gap before anyone treated the output as production grade. I also built multi‑step and agentic workflows with product teams, improving model performance on the task benchmarks the team tracked.
- Scale AI (2023 to 2024, Platinum Team Lead). I led a team of coding experts producing software training data, and audited training tasks in C++, Python, JavaScript, Java, and SQL against the project specification. The audit is where data quality is actually decided, so I read the spec as closely as the submission and sent back the cases where the two disagreed. I ran recruiting, onboarding, and interviewing as the program scaled, and directed a team of interviewers.
- My own products (ongoing). The worked example below is mine end to end, in repositories I can walk through line by line.
A prompt I built and measured
I use an LLM to draft and edit all the published copy for One Group Central, a member portal product I build. The copy was factually correct and read as machine written, which on a product whose pitch is trustworthy elections costs the sale.
My first attempt was wrong, and the way it was wrong is the useful part. I assumed a vocabulary problem, gave the model a banned word list, and watched it plateau. The tell is not the words. It is the habits, and habits survive a synonym swap: dashes as the default connector, lists shaped as bold label then colon then explanation, groups of three by reflex, a closing recap. The second thing I got wrong was worse. "Write more naturally" is unfalsifiable, so neither the model nor I could tell whether it had complied, and the behavior drifted back within a session or two.
So I rewrote the instruction as something checkable, moved it out of the chat and into the project's system prompt where every session and every subagent loads it, and then built the thing that makes it stick: a 239‑line checker that extracts prose from HTML, markdown, JavaScript, and Python, applies the rules per line, and exits nonzero on a violation.
- Measured on 2 October 2026, running the current checker against both corpora: the site copy as it stood before the rewrite carried 67 violations across 40 files. The current gated set is 0 across 118 files.
- The regression gate matters more than the sweep. Most of the site is generated, so hand‑fixing the HTML would have been undone by the next build. The generators hard fail on a violation, which means a regenerate cannot reintroduce what the sweep removed.
- The iteration I would defend hardest is the one where I removed rules. My first checker flagged "navigate", "unlock", and "highlight" as errors, and those words have correct literal uses. Failing a build on correct usage trained me to skim past my own gate, which is worse than having no gate, so those moved to a warning tier a person judges. A check with false positives does not fail loudly. It fails by being ignored.
- Two files still fail by design and sit on a named skip list: the privacy policy and the terms, because rewording a legal document changes what it promises.
What the number does not prove: the checker measures compliance with my rules, not quality. Zero errors means the known tells are gone, not that the writing is good, which is why the gate is mechanical plus a human reread. The validation I have not run is a blind test of before‑and‑after pairs against real readers, and that is the next thing I would build.
How I think about evaluation
- Golden set first, hand labeled, before touching the prompt. Hold out a slice. Decide the metric and the pass bar before seeing results, because deciding after is how a bar gets moved.
- Establish the measurement noise before trusting the gate. Two runs of identical code once measured a coverage metric at 44.28 and 44.27, and the second failed against the first. A gate finer than its own noise fails at random, and a flaky gate gets switched off. Non‑deterministic model evals have the same problem, and people respond by loosening the pass bar instead of widening the band.
- Coverage proves code ran. Only mutation testing proves an assertion would have failed if the behavior broke. The same question applies to an eval suite that passes without discriminating, and it is the first question I would ask of an LLM judge.
- Build tests as a confusion matrix, not as a list of things that work. My content classifier has 9 unit and 11 integration tests, and the ones I am proudest of assert the negatives: a member writing about harassment, assault, or being raped is never blocked, because those are the words somebody needs in order to report something.
- Say what a result does not prove. Every number on this page has a sentence attached about its limits, and that is deliberate.
Guardrails
The rules I actually care about are enforced outside the model, because a rule a model can be argued out of is a preference and a rule in the deploy pipeline is a constraint. Three that are in production:
- Content arriving through a tool is data and never an instruction, whatever it claims about its own authority.
- An allowlist in place of interpolation. The marketing contact form takes a slug that selects which of my own email bodies goes out, exact match, never echoed, so it cannot be used to mail arbitrary text.
- The server never fetches a user‑supplied URL. Having a backend retrieve an arbitrary address is a server‑side request forgery hole that reaches internal addresses with our credentials, so the client fetches and uploads the bytes instead.
What I haven't done
Worth stating plainly, because an engineer who will not tell you the edge of their experience is expensive to hire. I have not built a RAG pipeline or run a vector database in production. I have not implemented an LLM‑as‑judge, though I know how I would validate one. No fine tuning, no RLHF. No production multi‑agent framework, though I orchestrate subagents daily and have hit the real problem that comes with it, which is context isolation between them.
Let's talk
I'm in Mount Sinai, New York, and I work remotely. Available after 12pm Eastern on weekdays.