Listed on Intel’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role owns the communication layer that lets AI accelerator hardware clusters work as unified systems, designing collective algorithms and optimizing multi-node performance. It suits experienced systems programmers comfortable with low-level concurrency, hardware-software co-design, and the mathematical demands of distributed training at scale.
Our summary, not Intel’s wording. The full posting is on their site.
What they ask for
Required
- 5+ years AI/systems/HPC software experience in C++
- Strong concurrency and lock-free design
- Hands-on experience with collective libraries like NCCL or MPI
- Understanding of communication patterns in tensor, pipeline, and expert parallelism
- Performance engineering and profiling skills
Nice to have
- Topology and placement algorithms
- Congestion control experience
- Multi-tenant fabric experience
- Large-scale distributed training or inference exposure
- RDMA, InfiniBand, RoCE, or GPUDirect experience