Skill
Model Compression & Quantization
Data Science, Analytics and AI/ML
These are techniques used in machine learning to reduce the size and computational cost of trained neural network models, making them faster and more efficient to deploy, especially on mobile or edge devices. Quantization reduces the numerical precision of model weights (e.g., from 32-bit floats to 8-bit integers), while other compression methods include pruning and knowledge distillation. Machine learning engineers use these techniques to deploy large models like LLMs or vision models on hardware with limited memory and compute.
Model Compression & Quantization in the job market
Last checked September 12, 2026
- Open roles
- 15
- Employers
- 12
- Median pay
- —
- Disclose pay
- —
on Career App right now
hiring for it
not enough disclosed
of these roles
Open roles requiring Model Compression & Quantization (15)
Development Architect (f/m/d) - AI Foundation Model Training and Serving
SAP
Full-time · Potsdam, DE, 14469
Design and lead the AI infrastructure for a distributed platform enabling foundation model training and deployment across enterprises. This role suits architects with deep expertise in large-scale ML systems who can bridge research, product, and engineering teams while navigating complex federated environments.
Listed on SAP’s careers site · Apply there ↗
Employee Referral SRIB
Samsung Semiconductor
This is a referral program inviting you to recommend candidates with software engineering expertise across multiple domains including C++, AI/ML, cloud platforms, and hardware design. Samsung SRI-B is seeking talented engineers across various specializations from entry-level to experienced professionals.
Listed on Samsung Semiconductor’s careers site · Apply there ↗
Staff Machine Learning Engineer
Unity
Full-time · $218,400 – $283,900
This role involves optimizing state-of-the-art AI models to run efficiently on mobile and desktop devices within a browser-native runtime, handling everything from model export through kernel-level tuning to shipped features. It's ideal for a performance-focused engineer who thrives on closing the gap between research models and production on-device products, working with transformers, diffusion networks, and vision-language models across constrained hardware.
Listed on Unity’s careers site · Apply there ↗
AI Infrastructure Engineer
NIO
Full-time · San Jose, CA · $192,100 – $249,600
NIO is hiring a senior engineer to build production inference systems for large language and vision models across cloud and edge devices in their autonomous vehicle platform. This role suits someone with deep experience optimizing AI workloads on accelerators who wants to ship real-world impact at scale.
Listed on NIO’s careers site · Apply there ↗
Sr. Software Engineer
Ambarella
US Headquarters · $161,000 – $182,000
This role optimizes AI vision models for embedded processors, focusing on deploying deep learning onto specialized hardware like the Ambarella SoC. It suits engineers with expertise in PyTorch, computer vision, and embedded systems who want to bridge machine learning with hardware constraints.
Listed on Ambarella’s careers site · Apply there ↗
Embedded AI Engineer, On-Device Models
Deepgram
USA | Remote · $219,300 – $274,100
This role focuses on adapting Deepgram's speech AI models to run efficiently on resource-constrained devices like phones, wearables, and edge hardware. You'll optimize models for low power and memory through techniques like quantization and pruning, write performance-critical embedded code, and integrate with hardware accelerators to bring real-time voice AI to consumer devices.
Listed on Deepgram’s careers site · Apply there ↗
Member of Technical Staff - Research, Inference
Modal
New York, NY · $150,000 – $350,000
Modal is hiring a researcher to lead inference optimizations for their LLM serving platform, focusing on techniques like speculative decoding and quantization that reduce cost and latency. This role suits someone with a background shipping inference systems or research who can independently drive projects from conception through deployment.
Listed on Modal’s careers site · Apply there ↗
Senior Director, Developer Advocacy
Crusoe
Bellevue, WA · $280,000 – $300,000
This role is the public face of Crusoe Cloud to developers and technical decision makers, combining technical credibility with content creation and community engagement to build awareness and adoption. You'll create short-form video, define developer engagement strategy, and personally represent the company at events and across social platforms to technical audiences.
Listed on Crusoe’s careers site · Apply there ↗
Staff Machine Learning Engineer
XPeng
Full-time · Santa Clara, CA · $215,280 – $364,320
XPeng is seeking a machine learning engineer to develop and optimize traffic sign detection models for autonomous driving systems, handling the full lifecycle from data preparation through production deployment. This role suits experienced computer vision practitioners who want to work on safety-critical perception problems with real-world impact.
Listed on XPeng’s careers site · Apply there ↗
ML Engineer, Inference & Optimization
Pika Labs
Palo Alto, CA · $250,000 – $350,000
A senior or staff-level role optimizing AI model inference performance at a video generation startup. This suits engineers with deep expertise in GPU acceleration, distributed systems, and deploying large language models and video models efficiently at scale.
Listed on Pika Labs’s careers site · Apply there ↗
Asked for alongside Model Compression & Quantization
Measured from the 15 open roles that name Model Compression & Quantization — not from a curated list.
- C++8 roles53.3%
- CUDA8 roles53.3%
- Python7 roles46.7%
- PyTorch5 roles33.3%
- TensorFlow3 roles20.0%
Roles that use Model Compression & Quantization
Employers hiring for Model Compression & Quantization
12 in total, most open roles first.
Related skills
Curated neighbors in the taxonomy, whether or not employers ask for them together.