# Research Engineer, ML Infrastructure

Hiring organization: [Cognition AI](https://career.thegoodapps.co/organizations/cognition-ai)

Canonical page: https://career.thegoodapps.co/jobs/53ce8346-24ef-4e8c-a0eb-280d133fdb5f

Listed on Cognition AI's own careers site. Applications go to them directly.

- Location: San Francisco, CA

## Summary

This role focuses on building the distributed systems and infrastructure that enable AI research at scale, owning everything from GPU cluster management to experiment orchestration. It suits experienced systems engineers who understand both distributed computing fundamentals and deep learning frameworks deeply enough to anticipate researcher needs and eliminate bottlenecks.

_Our summary, not Cognition AI's wording._

## Skills named

C++, Python, PyTorch

## Required

- Deep experience building and operating distributed training systems for large models
- Strong systems engineering fundamentals including distributed systems, networking, and storage
- Proficiency in Python and C++
- Experience with PyTorch or equivalent deep learning framework at systems level
- Hands-on GPU performance profiling and memory optimization
- Experience implementing or optimizing parallelism strategies for large model training
- Strong debugging instincts for complex distributed systems
- ML knowledge sufficient to engage substantively with researchers

## Nice to have

- PhD or equivalent advanced credential

Apply on Cognition AI's site: https://jobs.ashbyhq.com/cognition/b6f96827-ce14-44f4-98fa-b1b8640858b6/application
