# Machine Learning Engineer - Inference

Hiring organization: [Together AI](https://career.thegoodapps.co/organizations/together-ai)

Canonical page: https://career.thegoodapps.co/jobs/295b7b41-88e0-4ac0-b316-8bc90cca4939

Listed on Together AI's own careers site. Applications go to them directly.

- Employment type: full time
- Seniority: Mid
- Location: San Francisco, CA
- Salary: 160000 – 230000 USD per year

## Summary

A Machine Learning Engineer role focused on building and optimizing inference systems for large language models at scale. This suits engineers with strong systems programming skills who want to work on high-performance AI infrastructure and collaborate closely with researchers.

_Our summary, not Together AI's wording._

## Skills named

CUDA, Python, PyTorch, Rust

## Required

- 3+ years building production-quality, high-performance code
- Proficiency with Python and PyTorch
- Experience building high-performance libraries and tooling
- Strong understanding of operating systems concepts including multi-threading, memory management, networking, and performance at scale

## Nice to have

- Knowledge of AI inference frameworks like TGI, vLLM, TensorRT-LLM, or Optimum
- Knowledge of inference optimization techniques such as speculative decoding
- CUDA or Triton programming experience
- Familiarity with Rust, Cython, or compilers

Apply on Together AI's site: https://job-boards.greenhouse.io/togetherai/jobs/4385540007
