# Forward Deployed Engineer (Inference & Post-Training)

Hiring organization: [Together AI](https://career.thegoodapps.co/organizations/together-ai)

Canonical page: https://career.thegoodapps.co/jobs/f7be2738-0493-4a32-9c10-e0fa30046b56

Listed on Together AI's own careers site. Applications go to them directly.

- Employment type: full time
- Seniority: Senior
- Location: San Francisco, CA
- Salary: 270000 – 300000 USD per year

## Summary

A technical specialist working directly with large customers to optimize AI model inference and fine-tuning on production systems. This role suits engineers with deep expertise in LLM deployment and performance tuning who want to work closely with strategic customers while influencing product direction.

_Our summary, not Together AI's wording._

## Skills named

Model Compression & Quantization, Python

## Required

- 5+ years in a technical role
- Expert-level hands-on experience with inference engines like vLLM, TensorRT-LLM, or SGLang
- Deep knowledge of KV cache tuning, speculative decoding, tensor parallelism, and quantization
- Hands-on experience with fine-tuning pipelines including LoRA, SFT, DPO, RLHF, and GRPO
- Strong Python skills
- Broad knowledge of open-source LLM landscape

Apply on Together AI's site: https://job-boards.greenhouse.io/togetherai/jobs/5131941007
