$155,000 – $245,000
Listed on Deepgram’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role adapts Deepgram's speech AI models to run on non-GPU edge hardware and embedded platforms, handling quantization, operator substitution, and deployment automation to get production models onto customer devices with minimal changes. It suits an experienced engineer who has shipped models to constrained hardware and wants to own that process across many platforms.
Our summary, not Deepgram’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Production experience deploying ML models to edge or non-NVIDIA hardware
- Working knowledge of quantization and precision tradeoffs
- Experience with at least one edge inference runtime and conversion toolchain
- Ability to modify and rewrite model graphs and swap unsupported operators
- Strong Python and PyTorch with production engineering practices
- Comfort building automation around model conversion and deployment
- Builder mindset with clear communication and ability to scope unfamiliar platforms
Nice to have
- Experience with speech, audio, or streaming/real-time models
- Ability to read or write low-level kernels
- Experience with model security or integrity on deployed devices
- Familiarity with multiple accelerator families
- Track record of building internal tooling for faster model porting or deployment