# Senior Software Engineer, Observability

Hiring organization: [Together AI](https://career.thegoodapps.co/organizations/together-ai)

Canonical page: https://career.thegoodapps.co/jobs/0397aba8-5394-4feb-a9c2-514fc4810d5e

Listed on Together AI's own careers site. Applications go to them directly.

- Employment type: full time
- Seniority: Senior
- Location: San Francisco, CA
- Salary: 200000 – 280000 USD per year

## Summary

This role leads the design and operation of observability platforms for a generative AI infrastructure company, managing metrics, logs, traces, and monitoring systems at scale. It suits engineers with deep expertise in distributed systems, cloud-native monitoring tools, and infrastructure automation who want to solve foundational challenges in AI infrastructure.

_Our summary, not Together AI's wording._

## Skills named

Ansible, ClickHouse, Docker, Go, Grafana, Helm, Kubernetes, MongoDB, OpenTelemetry, PostgreSQL, Prometheus, Python, Redis, Terraform

## Required

- Expertise in observability platforms (Prometheus, Grafana, ClickStack, OpenTelemetry)
- Experience with cloud-native monitoring (AWS, GCP, Azure)
- Strong programming in Go, Python, or equivalent
- Proficiency in infrastructure-as-code tools (Terraform, Ansible, Helm)
- Experience designing and scaling large-scale distributed systems
- Knowledge of containerization (Docker) and orchestration (Kubernetes)
- Understanding of microservices architecture and service mesh technologies
- Experience with CI/CD pipelines and GitOps workflows
- Knowledge of databases (PostgreSQL, MongoDB, Redis) and time-series databases

## Nice to have

- Monitoring AI/ML infrastructure and GPU clusters
- High-frequency, low-latency systems monitoring experience
- Chaos engineering and reliability testing background
- Open-source observability contributions
- Security monitoring and compliance framework knowledge

Apply on Together AI's site: https://job-boards.greenhouse.io/togetherai/jobs/4971774007
