$200,000 – $280,000
Listed on Together AI’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role leads the design and operation of observability platforms for a generative AI infrastructure company, managing metrics, logs, traces, and monitoring systems at scale. It suits engineers with deep expertise in distributed systems, cloud-native monitoring tools, and infrastructure automation who want to solve foundational challenges in AI infrastructure.
Our summary, not Together AI’s wording. The full posting is on their site.
Skills this role names
- Ansible
- ClickHouse
- Docker
- Go
- Grafana
- Helm
- Kubernetes
- MongoDB
- OpenTelemetry
- PostgreSQL
- Prometheus
- Python
- Redis
- Terraform
Log in to see which of these are already on your profile.
What they ask for
Required
- Expertise in observability platforms (Prometheus, Grafana, ClickStack, OpenTelemetry)
- Experience with cloud-native monitoring (AWS, GCP, Azure)
- Strong programming in Go, Python, or equivalent
- Proficiency in infrastructure-as-code tools (Terraform, Ansible, Helm)
- Experience designing and scaling large-scale distributed systems
- Knowledge of containerization (Docker) and orchestration (Kubernetes)
- Understanding of microservices architecture and service mesh technologies
- Experience with CI/CD pipelines and GitOps workflows
- Knowledge of databases (PostgreSQL, MongoDB, Redis) and time-series databases
Nice to have
- Monitoring AI/ML infrastructure and GPU clusters
- High-frequency, low-latency systems monitoring experience
- Chaos engineering and reliability testing background
- Open-source observability contributions
- Security monitoring and compliance framework knowledge