# Machine Learning Infrastructure Engineer

Hiring organization: [CharacterAI](https://career.thegoodapps.co/organizations/characterai)

Canonical page: https://career.thegoodapps.co/jobs/eb635ecb-6b8c-46d3-a845-f9a807893dfc

Listed on CharacterAI's own careers site. Applications go to them directly.

- Seniority: Senior
- Location: Redwood City, CA
- Remote: yes
- Salary: 150000 – 350000 USD per year

## Summary

This role supports machine learning research and products by building and maintaining GPU infrastructure, cluster diagnostics tools, and experiment management systems. It suits engineers with deep experience in ML operations who want to optimize hardware utilization and solve large-scale training and serving challenges.

_Our summary, not CharacterAI's wording._

## Skills named

Cloud Storage, JAX, Kubernetes, PyTorch, TensorFlow

## Required

- 4+ years supporting ML infrastructure
- Experience diagnosing ML infrastructure problems and failures
- Cloud platform experience (Compute Engine, Kubernetes, Cloud Storage)
- GPU experience

## Nice to have

- Large GPU cluster experience
- High-performance computing and networking experience
- Large language model training experience
- GPU kernel development experience

Apply on CharacterAI's site: https://jobs.ashbyhq.com/character/13006848-6a69-4467-9274-5d54e80b965d/application
