Skip to main content
CareerApp

Research Engineer, ML Infrastructure

Cognition AI

San Francisco, CA

Listed on Cognition AI’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

This role focuses on building the distributed systems and infrastructure that enable AI research at scale, owning everything from GPU cluster management to experiment orchestration. It suits experienced systems engineers who understand both distributed computing fundamentals and deep learning frameworks deeply enough to anticipate researcher needs and eliminate bottlenecks.

Our summary, not Cognition AI’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • Deep experience building and operating distributed training systems for large models
  • Strong systems engineering fundamentals including distributed systems, networking, and storage
  • Proficiency in Python and C++
  • Experience with PyTorch or equivalent deep learning framework at systems level
  • Hands-on GPU performance profiling and memory optimization
  • Experience implementing or optimizing parallelism strategies for large model training
  • Strong debugging instincts for complex distributed systems
  • ML knowledge sufficient to engage substantively with researchers

Nice to have

  • PhD or equivalent advanced credential

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.