$160,000 – $230,000
Listed on Together AI’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
A technical support engineer who handles customer issues with AI inference and fine-tuning platforms, working weekend shifts after an initial ramp-up period. The role bridges customer needs and internal engineering teams, requiring deep infrastructure expertise and hands-on troubleshooting of GPU clusters and Kubernetes deployments.
Our summary, not Together AI’s wording. The full posting is on their site.
Skills this role names
- Amazon Web Services (AWS)
- Ansible
- Git
- Grafana
- Infrastructure as Code
- JavaScript
- Kubernetes
- Microsoft Azure
- Postman
- Prometheus
- Python
- REST API
- TypeScript
Log in to see which of these are already on your profile.
What they ask for
Required
- 6+ years in customer-facing technical, SRE, DevOps, or infrastructure engineering role
- 1+ year supporting an AI service
- SRE or DevOps experience with Kubernetes
- Knowledge of AI, ML, GPU technologies, and HPC integration
- Production-level experience with Kubernetes, SLURM, and infrastructure-as-code tools
- Network diagnostic and trace analysis skills
- Python, TypeScript, or JavaScript proficiency
- Experience with observability tools like Prometheus or Grafana
- REST API debugging and HTTP knowledge
- LLM inference framework experience
- GPU cluster management background
- Cloud platform experience on AWS, GCP, or Azure
Nice to have
- Storage system operation experience in HPC environments (Vast, Weka)
- Hardware and platform migration expertise
- Complex problem-solving in dynamic environments