Skip to main content
CareerApp

Product Designer, Evals & Prompts

Menlo Ventures Portfolio

San Francisco, CA · full time

$305,000 – $385,000

Listed on Menlo Ventures Portfolio’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

This role builds the evaluation systems that test whether Claude's prompts and features work as intended across product surfaces and model launches. It combines eval infrastructure, prompt engineering support, and internal tooling to help product designers systematically test and refine AI behavior without writing code.

Our summary, not Menlo Ventures Portfolio’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • Production-quality Python code
  • Experience building evaluation pipelines for LLM products including graders and rubrics
  • Experience building internal tools with UI for non-technical users
  • Experience setting up test harnesses and sandboxing tool calls
  • Experience shipping prompts or working closely with prompt engineers
  • Ability to read and analyze transcripts, not just interpret scores

Nice to have

  • Worked inside a model-launch cycle
  • A/B testing experience and connecting offline evals to online results
  • Front-end or notebook-to-app experience with opinions on eval result visualization
  • Experience converting product rubrics into training signal for model improvement
  • Focus on how Claude behaves for users, beyond metric optimization

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.