Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.
52 open listings · Page 1 of 3
Roles posted by organizations here, alongside roles we found on employers’ own careers sites. Crawled roles say so on the card and send you to the employer to apply. Only show organizations on Career App
American International Group
Atlanta, GA
AIG is building a dedicated generative AI team and seeking an engineer to establish monitoring and governance systems that keep AI models performing fairly and reliably in production. This role suits someone who combines technical depth in machine learning and data engineering with a passion for responsible AI practices and can bridge the gap between technical teams and compliance frameworks.
Listed on American International Group’s careers site · Apply there ↗
SAP
Full-time · São Leopoldo, BR, 93022-718
This role builds platform infrastructure for secure data connectivity and access within SAP's Big Data Fabric Services, requiring deep expertise in cloud-native architecture, distributed systems security, and enterprise data integration. It suits experienced engineers who thrive on platform challenges spanning security, scalability, and operational excellence across hybrid SAP landscapes.
Listed on SAP’s careers site · Apply there ↗
Moderna
This role involves building and leading an enterprise-wide observability platform that monitors applications, infrastructure, and AI systems across Moderna's global operations. The position suits experienced engineers who want to shape platform strategy, work with modern technologies like OpenTelemetry and Grafana, and drive operational excellence in a regulated healthcare environment.
Listed on Moderna’s careers site · Apply there ↗
Target
7000 Target Pkwy N,NCD-0375 Brooklyn Park,MN 55445 · $98,000 – $176,000
This role combines AI engineering with backend services development, building intelligent systems for Target's financial platform. It suits engineers comfortable designing production AI applications alongside distributed systems who want to influence architecture and mentor others.
Listed on Target’s careers site · Apply there ↗
Workday
Full-time · Reston, VA · $163,800 – $245,800
This role involves building observability and logging infrastructure for Workday's developer platform, serving U.S. federal agencies. You'll design backend services, APIs, and tooling that help developers understand system performance and reliability in highly secure, mission-critical environments.
Listed on Workday’s careers site · Apply there ↗
NVIDIA
Full-time · Santa Clara, CA · $224,000 – $431,250
Lead a team building an agentic platform to monitor, debug, and optimize large-scale generative AI models in production. This role suits experienced engineering managers who understand deep learning systems, observability platforms, and distributed inference infrastructure.
Listed on NVIDIA’s careers site · Apply there ↗
LSEG
Saint Louis, MO
This role owns the CI/CD pipelines and infrastructure automation for LSEG's internal observability platform, spanning from code commit through production deployment across multiple cloud environments. It suits engineers who thrive on full-stack delivery ownership, from GitOps workflows and Terraform infrastructure to Kubernetes operations, with the chance to build tooling that directly impacts how a strategic platform gets deployed globally.
Listed on LSEG’s careers site · Apply there ↗
CrowdStrike
Full-time · Sunnyvale, CA · $140,000 – $215,000
Lead a team of Site Reliability Engineers at a major cybersecurity firm, responsible for managing infrastructure that processes trillions of daily events while mentoring engineers and driving operational excellence across distributed systems. This role suits experienced engineering managers with deep SRE expertise who want to balance hands-on technical leadership with team growth at scale.
Listed on CrowdStrike’s careers site · Apply there ↗
Comcast
Full-time · $143,911 – $215,866
Lead a team building automation infrastructure for Comcast's nationwide content delivery network, managing engineers across distributed locations while overseeing the technical direction and operational excellence of large-scale CDN systems. This role suits experienced engineering managers with deep infrastructure expertise who can balance hands-on technical decision-making with people leadership and stakeholder communication.
Listed on Comcast’s careers site · Apply there ↗
WP Engine
PLN 244,000 – PLN 335,500
A senior backend engineer role focused on designing and operating scalable cloud infrastructure services for a WordPress platform company serving millions of customers globally. This position suits engineers with strong fundamentals in distributed systems and production reliability who want full ownership of services across their entire lifecycle.
Listed on WP Engine’s careers site · Apply there ↗
Grafana Labs
Full-time · United States (Remote) · $154,445 – $185,334
A full-stack engineering role on Grafana Cloud's observability integrations team, building features across React frontends, Go backends, and Jsonnet configurations. Suited for someone comfortable shipping complete features from design through production and representing work cross-functionally while contributing to open-source observability tools.
Listed on Grafana Labs’s careers site · Apply there ↗
Kong Company
Washington, United States · $153,000 – $218,000 · Remote
Lead a newly-formed team of Site Reliability Engineers managing Kong's cloud-hosted API gateway service, responsible for maintaining enterprise-grade uptime and performance while staying hands-on with critical implementations. The role suits experienced engineering managers who thrive in high-ownership environments and can balance strategic team-building with direct technical contribution.
Amplitude
Full-time · San Francisco, CA · $227,000 – $381,000
Lead Amplitude's Data Infrastructure team within their experimentation platform, connecting statistical innovation with scalable production systems that power large-scale A/B testing. This role combines technical leadership with data science expertise, requiring someone equally fluent in causal inference methods and distributed computing architecture.
Listed on Amplitude’s careers site · Apply there ↗
Honeycomb
Full-time · Remote - United States · $183,340 – $206,000
This role involves designing and maintaining backend systems and APIs in Go while contributing full-stack features in React and TypeScript for Honeycomb's AI observability platform. It suits experienced backend engineers who enjoy technical leadership, mentoring, and collaborating across product teams to build observability infrastructure for AI workloads.
Listed on Honeycomb’s careers site · Apply there ↗
Render
Remote: United States · $218,000 – $300,000
This Staff Product Manager role owns the observability strategy for Render's platform, spanning traditional systems monitoring and the emerging domain of AI agent observability. It suits seasoned product leaders who combine deep infrastructure knowledge with hands-on familiarity of OpenTelemetry and LLM economics, and who can navigate the complex technical and business tradeoffs required to make AI workloads debuggable and cost-transparent.
Listed on Render’s careers site · Apply there ↗
Crusoe
Full-time · San Francisco, CA · $215,000 – $260,000
Lead a team of 4-6 engineers building Crusoe's telemetry agent, the observability software that collects metrics and logs across their AI cloud infrastructure. This role combines hands-on technical leadership with people management, shipping incremental releases against a fixed deadline while growing the team and expanding into adjacent areas like hardware validation.
Listed on Crusoe’s careers site · Apply there ↗
BlackSky
Full-time · Herndon, VA · $135,000 – $150,000 · Remote
BlackSky seeks an engineer to design and operate cloud infrastructure, Kubernetes clusters, and GitOps pipelines across public, private, and air-gapped networks for intelligence platforms. This role suits someone with deep Kubernetes expertise and hands-on experience managing enterprise-scale deployments in restrictive network environments.
Listed on BlackSky’s careers site · Apply there ↗
Supabase
Full-time · Remote, Global
A distributed SRE role focused on establishing reliability practices and frameworks across Supabase's engineering organization rather than owning infrastructure directly. This suits someone with deep SRE experience who wants to drive systemic improvements and influence multiple teams through expertise and collaboration.
Listed on Supabase’s careers site · Apply there ↗
InstaLily
Full-time · New York, NY · $150,000 – $190,000
InstaLILY is hiring a Site Reliability Engineer to build and operate their Internal Developer Platform and multi-cloud Kubernetes infrastructure that supports their AI agent products. You'll work in a small team with direct impact, partnering with engineers to create self-service tooling and golden paths while managing production systems that power major enterprise customers.
Listed on InstaLily’s careers site · Apply there ↗
Bend Studio
Full-time · San Diego, CA · $150,000 – $225,000
Sony Interactive Entertainment seeks an AI software engineer to design and deploy production AI systems—including LLM-powered workflows, retrieval-augmented generation, and agentic systems—that solve real business problems across PlayStation's digital commerce and player-facing services. This role suits a strong backend or platform engineer with hands-on experience building applied AI features and shipping them reliably to scale.
Listed on Bend Studio’s careers site · Apply there ↗
LaunchDarkly
2 locations · $145,500 – $235,400 · Remote
LaunchDarkly is hiring a full-stack engineer to build enterprise capabilities for their Observability product, focusing on infrastructure-as-code support and integrations with third-party tools. The role suits experienced engineers comfortable shipping across the full stack who want to work on features that directly impact enterprise adoption.
Listed on LaunchDarkly’s careers site · Apply there ↗
San Francisco, CA
Reddit is seeking a Staff-level Site Reliability Engineer to lead reliability efforts across their platform's user-facing systems at massive scale. This role combines technical depth in distributed systems with leadership responsibilities, focusing on improving availability, performance, and operational excellence across web, mobile, and real-time infrastructure.
Listed on Reddit’s careers site · Apply there ↗
LangChain
New York, NY · $130,000 – $195,000 · Remote
A senior technical support role focused on helping AI and infrastructure engineers troubleshoot production issues with LangChain's platform and open-source tools. This suits experienced technical support professionals who want to work with complex distributed systems and take on escalations while helping define support excellence in the AI space.
Listed on LangChain’s careers site · Apply there ↗
XPeng
Full-time · Santa Clara, CA · $244,140 – $413,160
XPENG is seeking a hands-on Senior Staff AI Engineer to design and deploy production-grade AI systems that leverage large language models, retrieval-augmented generation, and agentic workflows across the organization. The role suits experienced AI systems engineers who can architect reusable patterns, establish best practices, and drive AI adoption while partnering with product and engineering teams to deliver measurable impact.
Listed on XPeng’s careers site · Apply there ↗
Future Co
Full-time · Remote · $215,000 – $250,000
This role involves building and deploying AI agents and LLM-powered features that directly serve users in a health guidance platform, from conception through production monitoring. It suits engineers who enjoy shipping practical AI systems end-to-end and are comfortable working across the full stack from prompt engineering to infrastructure.
Listed on Future Co’s careers site · Apply there ↗