Skip to main content
CareerApp

Staff ML Engineer, Agent Training & Environments

Labelbox

San Francisco, CA · full time · Lead / Staff / Principal

$250,000 – $280,000

Listed on Labelbox’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

Build infrastructure and systems that enable frontier AI labs to train and evaluate agentic models, bridging high-throughput platform engineering with deep reinforcement learning expertise. This role combines designing RL environments and reward systems, implementing verification and grading pipelines, and scaling fine-tuning infrastructure for teams pushing the boundaries of AI agents.

Our summary, not Labelbox’s wording. The full posting is on their site.

What they ask for

Required

  • 3+ years shipping production systems others rely on
  • Strong system and API design judgment
  • Daily production code shipped with coding agents
  • Build infrastructure for team tooling, CI, and harnesses
  • Work effectively in ambiguous startup environments
  • Deep Python proficiency
  • Fine-tuned models for agentic tasks using SFT and at least one RL method
  • Built environments for agents to operate in
  • Designed verifiers or graders for open-ended work
  • Forensic debugging of training runs
  • Understand compute economics and experimental efficiency
  • Document and communicate learnings to the team

Nice to have

  • Experience with agent harnesses and coding agents in training and evaluation
  • Multi-tenancy and sandboxing for untrusted agent execution
  • Production distributed systems, ML infrastructure, or data systems at scale
  • Direct experience with frontier labs or highly technical customers

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.