$500,000 – $850,000
Listed on Menlo Ventures Portfolio’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role combines research and engineering to advance Claude's ability to write and debug code through reinforcement learning. You'll design RL environments, build reward systems, run training experiments, and optimize the infrastructure that powers these systems, working across agentic coding behaviors, correctness, and performance optimization.
Our summary, not Menlo Ventures Portfolio’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Strong software engineering skills and deep Python expertise, including async/concurrent programming
- Ability to own systems end to end and debug across the stack
- Balance research exploration with engineering implementation and rigorously shape experimental design and interpret results
- Care about code quality, testing, and performance
- Passion for AI safety and beneficial systems development
- Bachelor's degree or equivalent education, training, and professional experience
Nice to have
- Experience with reinforcement learning, RLHF, post-training, or LLM finetuning
- Built coding agents, code-execution sandboxes, eval harnesses, verifiers, or developer tooling
- Background in program analysis, testing, verification, compilers, or formal methods
- Experience with PyTorch and large-scale distributed training
- Performance profiling and optimization of ML systems
- CUDA/GPU or TPU kernel experience and accelerator-performance intuition
- Experience with virtualization and sandboxed code execution environments