# Research Engineer, Post-Training Inference

Hiring organization: [Together AI](https://career.thegoodapps.co/organizations/together-ai)

Canonical page: https://career.thegoodapps.co/jobs/e885cc10-a0a6-4d8f-865b-ea5811e0d5bb

Listed on Together AI's own careers site. Applications go to them directly.

- Employment type: full time
- Location: San Francisco, CA
- Salary: 200000 – 290000 USD per year

## Summary

This role involves building and optimizing systems that allow developers to fine-tune and deploy open-source AI models efficiently, working across the full pipeline from training through production inference. You'll collaborate on core infrastructure for customization services, focusing on inference optimization and integration between post-training and serving platforms.

_Our summary, not Together AI's wording._

## Skills named

CUDA, Go, Kubernetes, Python

## Required

- 2+ years building and deploying ML-based services in production
- Hands-on experience with modern inference engines such as SGLang, vLLM, or TensorRT-LLM
- Familiarity with modern fine-tuning methods for LLMs and AI models
- Strong software engineering background in Python or Go
- Awareness of current advances and trends in machine learning

## Nice to have

- Experience serving low-precision models or managing multiple LoRA adapters in one instance
- Experience optimizing RL training workloads
- Experience developing CUDA, Triton, or CuTE DSL kernels for inference
- Experience building large-scale, high-load production systems
- Contributions to or maintenance of open-source ML projects
- Experience managing ML workloads on Kubernetes

Apply on Together AI's site: https://job-boards.greenhouse.io/togetherai/jobs/5179372007
