Skip to main content
CareerApp

Performance Engineer, Inference Engine

Menlo Ventures Portfolio

New York, NY · Mid

$350,000 – $850,000

Listed on Menlo Ventures Portfolio’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

This role optimizes Anthropic's inference engine—the software layer that manages how their Claude language model runs on accelerators and serves requests efficiently. It's ideal for systems engineers who enjoy performance optimization work across hardware, networking, and distributed systems.

Our summary, not Menlo Ventures Portfolio’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • Understanding of LLM inference mechanics (prefill, decode, accelerator compute and memory)
  • Systems programming proficiency in Rust, C++, or equivalent languages
  • Performance analysis and profiling methodology
  • Strong code quality and testing practices
  • Ability to learn complex unfamiliar systems quickly and ship changes

Nice to have

  • Experience optimizing an LLM serving engine
  • GPU or accelerator programming
  • Operating system internals
  • Transformer-based language modeling
  • Experience building allocators, caches, schedulers, or high-bandwidth transports
  • Rust fluency
  • Systems reproducibility work (determinism, replay, property-based testing)

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.