# Performance Engineer, Inference Engine

Hiring organization: [Menlo Ventures Portfolio](https://career.thegoodapps.co/organizations/menlo-ventures-portfolio)

Canonical page: https://career.thegoodapps.co/jobs/9fdc8e5c-5fe9-4e3f-a42f-e8d67ef61d02

Listed on Menlo Ventures Portfolio's own careers site. Applications go to them directly.

- Seniority: Mid
- Location: New York, NY
- Salary: 350000 – 850000 USD per year

## Summary

This role optimizes Anthropic's inference engine—the software layer that manages how their Claude language model runs on accelerators and serves requests efficiently. It's ideal for systems engineers who enjoy performance optimization work across hardware, networking, and distributed systems.

_Our summary, not Menlo Ventures Portfolio's wording._

## Skills named

C++, CUDA, Rust

## Required

- Understanding of LLM inference mechanics (prefill, decode, accelerator compute and memory)
- Systems programming proficiency in Rust, C++, or equivalent languages
- Performance analysis and profiling methodology
- Strong code quality and testing practices
- Ability to learn complex unfamiliar systems quickly and ship changes

## Nice to have

- Experience optimizing an LLM serving engine
- GPU or accelerator programming
- Operating system internals
- Transformer-based language modeling
- Experience building allocators, caches, schedulers, or high-bandwidth transports
- Rust fluency
- Systems reproducibility work (determinism, replay, property-based testing)

Apply on Menlo Ventures Portfolio's site: https://job-boards.greenhouse.io/anthropic/jobs/5418323008
