Skip to main content
CareerApp

Senior Software Engineer, Observability

Together AI

San Francisco, CA · full time · Senior

$200,000 – $280,000

Listed on Together AI’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

This role leads the design and operation of observability platforms for a generative AI infrastructure company, managing metrics, logs, traces, and monitoring systems at scale. It suits engineers with deep expertise in distributed systems, cloud-native monitoring tools, and infrastructure automation who want to solve foundational challenges in AI infrastructure.

Our summary, not Together AI’s wording. The full posting is on their site.

What they ask for

Required

  • Expertise in observability platforms (Prometheus, Grafana, ClickStack, OpenTelemetry)
  • Experience with cloud-native monitoring (AWS, GCP, Azure)
  • Strong programming in Go, Python, or equivalent
  • Proficiency in infrastructure-as-code tools (Terraform, Ansible, Helm)
  • Experience designing and scaling large-scale distributed systems
  • Knowledge of containerization (Docker) and orchestration (Kubernetes)
  • Understanding of microservices architecture and service mesh technologies
  • Experience with CI/CD pipelines and GitOps workflows
  • Knowledge of databases (PostgreSQL, MongoDB, Redis) and time-series databases

Nice to have

  • Monitoring AI/ML infrastructure and GPU clusters
  • High-frequency, low-latency systems monitoring experience
  • Chaos engineering and reliability testing background
  • Open-source observability contributions
  • Security monitoring and compliance framework knowledge

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.