$219,300 – $274,100
Listed on Deepgram’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role focuses on adapting Deepgram's speech AI models to run efficiently on resource-constrained devices like phones, wearables, and edge hardware. You'll optimize models for low power and memory through techniques like quantization and pruning, write performance-critical embedded code, and integrate with hardware accelerators to bring real-time voice AI to consumer devices.
Our summary, not Deepgram’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Production experience on resource-constrained embedded systems, mobile, or edge AI
- Strong C, C++, or Rust proficiency for performance-critical constrained code
- Hands-on model optimization for on-device deployment including quantization or pruning
- Familiarity with edge inference runtimes like ONNX Runtime, TensorRT, or TFLite
- Understanding of CPU, GPU, NPU, DSP architectures and memory hierarchies
- Experience in bare-metal or RTOS environments such as FreeRTOS or Zephyr
- Strong communication and ability to scope and drive optimization problems to measurable results
Nice to have
- Real-time audio processing on embedded platforms or wake-word detection
- Depth in ML optimization techniques like mixed-precision inference or neural architecture search
- Hardware evaluation and benchmarking experience across accelerators and SoCs
- Shipped AI features in consumer products at scale
- Experience with model compilation toolchain tradeoffs across hardware targets
- Secure on-device deployment practices including code signing and encrypted model storage