$140,000 – $200,000
Listed on Labelbox’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role owns the design and operationalization of reinforcement learning training environments—the sandboxed systems where AI agents learn and interact. It's ideal for a systems-minded software engineer who understands RL fundamentals and wants to solve high-leverage infrastructure problems at the intersection of ML and DevOps.
Our summary, not Labelbox’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- 2+ years professional software engineering experience
- Strong Python fundamentals
- At least one systems-level language (Go, Rust, or C++)
- Production containerization and sandboxing experience (Docker, Podman, Firecracker or similar)
- Understanding of RL concepts: MDPs, reward shaping, episode structure, observation/action spaces
- Experience building or maintaining developer tooling, CLI tools, or infrastructure automation
- Comfort with browser automation or terminal interaction tooling
- Strong debugging across process boundaries and container layers
- Ability to implement from academic papers and open-source benchmarks independently
Nice to have
- Direct experience building RL environments (Gymnasium/Gym, PettingZoo, or custom implementations)
- Experience with agentic AI evaluation frameworks (SWE-bench, WebArena, OSWorld, TerminalBench)
- GCP or AWS infrastructure experience
- Prior work at AI data, ML platform, or AI research lab
- Open-source contributions in RL, agents, or dev-tools