Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.
106 open listings · Page 2 of 5
Roles posted by organizations here, alongside roles we found on employers’ own careers sites. Crawled roles say so on the card and send you to the employer to apply. Only show organizations on Career App
Intel
Full-time · $195,200 – $361,200
This role leads the development of firmware and runtime software for Intel's neuromorphic AI accelerators, targeting edge and robotic systems. It suits experienced systems software architects who combine deep expertise in low-level performance optimization with the ability to guide technical direction across hardware-software integration challenges.
Listed on Intel’s careers site · Apply there ↗
Unity
Full-time · 2 locations · $218,400 – $283,900
This role involves optimizing state-of-the-art AI models to run efficiently on mobile and desktop devices within a browser-native runtime, handling everything from model export through kernel-level tuning to shipped features. It's ideal for a performance-focused engineer who thrives on closing the gap between research models and production on-device products, working with transformers, diffusion networks, and vision-language models across constrained hardware.
Listed on Unity’s careers site · Apply there ↗
Together AI
Full-time · San Francisco, CA · $240,000 – $280,000
This role owns the software infrastructure that automatically provisions and manages GPU clusters for AI training and inference, turning manual provisioning into a self-service API. You'll build state machines, orchestration systems, and self-healing automation that eliminate manual infrastructure work, built and operated like a production software platform.
Listed on Together AI’s careers site · Apply there ↗
Menlo Ventures Portfolio
Full-time · New York, NY · $500,000 – $850,000
This role combines research and engineering to advance Claude's ability to write and debug code through reinforcement learning. You'll design RL environments, build reward systems, run training experiments, and optimize the infrastructure that powers these systems, working across agentic coding behaviors, correctness, and performance optimization.
Listed on Menlo Ventures Portfolio’s careers site · Apply there ↗
Motional
Remote U.S. · $240,000 – $330,000
This role leads the machine learning team responsible for motion planning and behavior prediction on autonomous vehicles, working on models that help self-driving cars navigate traffic safely and predict other road users' actions. It suits experienced ML engineers and technical leaders with a background in autonomous systems who want to shape critical safety-focused technology.
Listed on Motional’s careers site · Apply there ↗
Crusoe
Full-time · San Francisco, CA · $200,000 – $240,000
This role leads cross-functional program delivery for Crusoe's managed LLM inference platform, coordinating model engineering, infrastructure, and operations through multi-quarter release cycles. It suits experienced technical program managers with deep knowledge of how large language models are served and optimized in production at scale.
Listed on Crusoe’s careers site · Apply there ↗
Modal
New York, NY · $200,000 – $350,000
Modal is seeking an engineer to optimize machine learning inference and fine-tuning workloads on their GPU infrastructure platform. This role suits someone with deep experience in ML systems performance who enjoys debugging GPU bottlenecks and shipping optimizations at scale.
Listed on Modal’s careers site · Apply there ↗
XPeng
Full-time · Santa Clara, CA · $215,280 – $364,320
XPENG is hiring a staff-level ML engineer to optimize large language models and foundation models for autonomous driving applications, focusing on inference efficiency and training acceleration. This role suits someone with deep expertise in transformer optimization, GPU programming, and model deployment who wants to work on cutting-edge AI infrastructure for self-driving vehicles.
Listed on XPeng’s careers site · Apply there ↗
Cohere
Full-time · New York, NY · CA$250,000 – CA$535,000 · Remote
Cohere seeks a senior engineer to build and maintain the training framework powering their large-scale language models, working across distributed systems, HPC infrastructure, and tooling. This role suits someone with deep expertise in distributed training systems who wants ownership over critical ML infrastructure components.
Listed on Cohere’s careers site · Apply there ↗
AeroVironment
Full-time · Huntsville, AL
This role designs and develops physics-based 3D synthetic environments for military simulation and testing, combining radiometry expertise with graphics programming. It suits experienced software engineers comfortable with C++, Linux, and rendering systems who want to work on specialized defense applications.
Listed on AeroVironment’s careers site · Apply there ↗
NVIDIA
$184,000 – $287,500
NVIDIA is seeking a senior C++ engineer to design and optimize foundational libraries and algorithms for their CUDA GPU programming platform. This role suits experienced systems programmers who enjoy balancing performance optimization, API design, and developer productivity across multiple language ecosystems.
Listed on NVIDIA’s careers site · Apply there ↗
Intel
Full-time · $170,500 – $315,490
Intel is seeking a performance engineer to optimize Large Language Model inference on their next-generation GPUs, working across the full stack from kernel development to open-source framework contributions. This role suits engineers passionate about squeezing maximum throughput from hardware and collaborating with the broader AI infrastructure community.
Listed on Intel’s careers site · Apply there ↗
Together AI
Full-time · San Francisco, CA · $200,000 – $290,000
This role involves building and optimizing systems that allow developers to fine-tune and deploy open-source AI models efficiently, working across the full pipeline from training through production inference. You'll collaborate on core infrastructure for customization services, focusing on inference optimization and integration between post-training and serving platforms.
Listed on Together AI’s careers site · Apply there ↗
Motional
Boston, MA · $144,000 – $192,000 · Remote
This role optimizes the systems and infrastructure that allow ML researchers to train large models efficiently, focusing on performance profiling, distributed training, and GPU kernel development. It suits engineers who combine deep systems knowledge with hands-on ML experience and want to work on the foundational tech that powers next-generation model training.
Listed on Motional’s careers site · Apply there ↗
Crusoe
Full-time · San Francisco, CA · $172,500 – $210,000
This role involves designing and maintaining automated testing frameworks for large-scale GPU clusters, with a focus on validating interconnect performance, distributed workload scaling, and multi-node stability in virtualized environments. It suits engineers with deep systems knowledge who enjoy working on low-level infrastructure challenges in a fast-moving AI compute company.
Listed on Crusoe’s careers site · Apply there ↗
Menlo Ventures Portfolio
Full-time · San Francisco, CA · $350,000 – $850,000
This role involves designing reinforcement learning environments and conducting research to improve AI models' code generation capabilities, particularly for accelerator programming. You'll need deep expertise in GPU/accelerator optimization and ML frameworks, combined with the ability to move research from experimentation into production training pipelines.
Listed on Menlo Ventures Portfolio’s careers site · Apply there ↗
XPeng
Full-time · Santa Clara, CA · $174,720 – $295,680
XPENG is seeking a machine learning engineer to optimize and deploy large language models for autonomous driving applications, focusing on inference efficiency across diverse hardware platforms. This role suits engineers with deep transformer knowledge and hands-on experience in kernel optimization, quantization, and model acceleration techniques.
Listed on XPeng’s careers site · Apply there ↗
Cohere
Full-time · New York, NY · CA$250,000 – CA$535,000 · Remote
This role involves optimizing how large language models execute in production environments, focusing on reducing latency and improving throughput. It's ideal for engineers who enjoy performance tuning, GPU optimization, and shipping measurable improvements to complex distributed systems.
Listed on Cohere’s careers site · Apply there ↗
NVIDIA
Santa Clara, CA · $272,000 – $431,250
A principal-level systems software engineer role focused on designing and implementing core GPU kernel scheduling features within NVIDIA's CUDA Driver platform. This position suits experienced C/C++ developers with deep OS-level knowledge who want to shape the software foundation enabling AI, scientific computing, and graphics workloads on NVIDIA hardware.
Listed on NVIDIA’s careers site · Apply there ↗
Together AI
Full-time · San Francisco, CA · $220,000 – $280,000
Staff-level machine learning engineer role focused on optimizing inference for voice models like speech-to-text and text-to-speech on Together AI's platform. Ideal for someone with deep expertise in LLM serving engines and GPU optimization who wants to shape the technical direction of a foundational voice AI system.
Listed on Together AI’s careers site · Apply there ↗
Motional
Boston, MA · $136,500 – $225,000 · Remote
This role develops computer vision and deep learning systems for autonomous vehicle perception, working on detection, classification, and segmentation tasks that feed into self-driving technology. It suits engineers with machine learning fundamentals and practical neural network experience who want to apply their skills to safety-critical, real-world autonomous driving applications.
Listed on Motional’s careers site · Apply there ↗
Menlo Ventures Portfolio
Full-time · San Francisco, CA · $315,000 – $560,000
This role builds the computational infrastructure and tooling that lets researchers understand how large language models actually work internally, rather than treating them as black boxes. It's suited to experienced software engineers who want to work on AI safety through hands-on infrastructure work, enjoy translating research needs into systems, and are comfortable optimizing across the full stack from GPU kernels to user-facing tools.
Listed on Menlo Ventures Portfolio’s careers site · Apply there ↗
Cohere
Full-time · New York, NY · CA$250,000 – CA$535,000 · Remote
This role optimizes training performance for Cohere's large language models, focusing on throughput and accelerator utilization through software engineering and low-level kernel optimization. It suits engineers with strong systems programming skills who want to work on cutting-edge AI infrastructure and training infrastructure at scale.
Listed on Cohere’s careers site · Apply there ↗
NVIDIA
2 locations · $152,000 – $287,500
NVIDIA is hiring a senior software engineer to optimize deep learning inference frameworks and GPU-accelerated pipelines for large language models and generative AI applications. This role suits experienced systems engineers who want to work on high-performance open-source tools and tackle complex performance optimization challenges across NVIDIA's accelerator hardware.
Listed on NVIDIA’s careers site · Apply there ↗
Together AI
Full-time · San Francisco, CA · $190,000 – $270,000
A systems engineering role focused on building and automating large-scale GPU infrastructure for AI model training and inference. This suits engineers who approach infrastructure as a software problem and are driven to eliminate manual operations through intelligent automation at massive scale.
Listed on Together AI’s careers site · Apply there ↗