Skip to main content
CareerApp

Skill

Model Compression & Quantization

Data Science, Analytics and AI/ML

These are techniques used in machine learning to reduce the size and computational cost of trained neural network models, making them faster and more efficient to deploy, especially on mobile or edge devices. Quantization reduces the numerical precision of model weights (e.g., from 32-bit floats to 8-bit integers), while other compression methods include pruning and knowledge distillation. Machine learning engineers use these techniques to deploy large models like LLMs or vision models on hardware with limited memory and compute.

See who is hiring

Model Compression & Quantization in the job market

Last checked September 12, 2026

Open roles
15

on Career App right now

Employers
12

hiring for it

Median pay

not enough disclosed

Disclose pay

of these roles

Open roles requiring Model Compression & Quantization (15)

Development Architect (f/m/d) - AI Foundation Model Training and Serving

SAP

Full-time · Potsdam, DE, 14469

Design and lead the AI infrastructure for a distributed platform enabling foundation model training and deployment across enterprises. This role suits architects with deep expertise in large-scale ML systems who can bridge research, product, and engineering teams while navigating complex federated environments.

Listed on SAP’s careers site · Apply there ↗

Samsung Semiconductor

Employee Referral SRIB

Samsung Semiconductor

This is a referral program inviting you to recommend candidates with software engineering expertise across multiple domains including C++, AI/ML, cloud platforms, and hardware design. Samsung SRI-B is seeking talented engineers across various specializations from entry-level to experienced professionals.

Listed on Samsung Semiconductor’s careers site · Apply there ↗

Staff Machine Learning Engineer

Unity

Full-time · $218,400 – $283,900

This role involves optimizing state-of-the-art AI models to run efficiently on mobile and desktop devices within a browser-native runtime, handling everything from model export through kernel-level tuning to shipped features. It's ideal for a performance-focused engineer who thrives on closing the gap between research models and production on-device products, working with transformers, diffusion networks, and vision-language models across constrained hardware.

Listed on Unity’s careers site · Apply there ↗

AI Infrastructure Engineer

NIO

Full-time · San Jose, CA · $192,100 – $249,600

NIO is hiring a senior engineer to build production inference systems for large language and vision models across cloud and edge devices in their autonomous vehicle platform. This role suits someone with deep experience optimizing AI workloads on accelerators who wants to ship real-world impact at scale.

Listed on NIO’s careers site · Apply there ↗

Sr. Software Engineer

Ambarella

US Headquarters · $161,000 – $182,000

This role optimizes AI vision models for embedded processors, focusing on deploying deep learning onto specialized hardware like the Ambarella SoC. It suits engineers with expertise in PyTorch, computer vision, and embedded systems who want to bridge machine learning with hardware constraints.

Listed on Ambarella’s careers site · Apply there ↗

Embedded AI Engineer, On-Device Models

Deepgram

USA | Remote · $219,300 – $274,100

This role focuses on adapting Deepgram's speech AI models to run efficiently on resource-constrained devices like phones, wearables, and edge hardware. You'll optimize models for low power and memory through techniques like quantization and pruning, write performance-critical embedded code, and integrate with hardware accelerators to bring real-time voice AI to consumer devices.

Listed on Deepgram’s careers site · Apply there ↗

Member of Technical Staff - Research, Inference

Modal

New York, NY · $150,000 – $350,000

Modal is hiring a researcher to lead inference optimizations for their LLM serving platform, focusing on techniques like speculative decoding and quantization that reduce cost and latency. This role suits someone with a background shipping inference systems or research who can independently drive projects from conception through deployment.

Listed on Modal’s careers site · Apply there ↗

Senior Director, Developer Advocacy

Crusoe

Bellevue, WA · $280,000 – $300,000

This role is the public face of Crusoe Cloud to developers and technical decision makers, combining technical credibility with content creation and community engagement to build awareness and adoption. You'll create short-form video, define developer engagement strategy, and personally represent the company at events and across social platforms to technical audiences.

Listed on Crusoe’s careers site · Apply there ↗

Apply for Staff Machine Learning Engineer at XPeng on their site

Staff Machine Learning Engineer

XPeng

Full-time · Santa Clara, CA · $215,280 – $364,320

XPeng is seeking a machine learning engineer to develop and optimize traffic sign detection models for autonomous driving systems, handling the full lifecycle from data preparation through production deployment. This role suits experienced computer vision practitioners who want to work on safety-critical perception problems with real-world impact.

Listed on XPeng’s careers site · Apply there ↗

ML Engineer, Inference & Optimization

Pika Labs

Palo Alto, CA · $250,000 – $350,000

A senior or staff-level role optimizing AI model inference performance at a video generation startup. This suits engineers with deep expertise in GPU acceleration, distributed systems, and deploying large language models and video models efficiently at scale.

Listed on Pika Labs’s careers site · Apply there ↗

View all 15 on Jobs

Asked for alongside Model Compression & Quantization

Measured from the 15 open roles that name Model Compression & Quantization — not from a curated list.

Employers hiring for Model Compression & Quantization

12 in total, most open roles first.

Related skills

Curated neighbors in the taxonomy, whether or not employers ask for them together.

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.