$218,400 – $283,900
Listed on Unity’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role involves optimizing state-of-the-art AI models to run efficiently on mobile and desktop devices within a browser-native runtime, handling everything from model export through kernel-level tuning to shipped features. It's ideal for a performance-focused engineer who thrives on closing the gap between research models and production on-device products, working with transformers, diffusion networks, and vision-language models across constrained hardware.
Our summary, not Unity’s wording. The full posting is on their site.
Skills this role names
- CUDA
- JavaScript
- Model Compression & Quantization
- Model Optimization
- Python
- Sheet Metal
- TypeScript
- Vulkan API
Log in to see which of these are already on your profile.
What they ask for
Required
- 5+ years in software or ML engineering with focus on on-device or edge inference
- Production deployment of transformer or diffusion models on mobile, desktop, or embedded hardware
- Hands-on experience with at least one major inference runtime (ONNX Runtime, CoreML, TFLite, or ExecuTorch)
- Low-level performance engineering with at least one GPU or compute API (WebGPU, Metal, Vulkan, D3D12, or CUDA)
- Working knowledge of model optimization techniques (quantization, weight sharing, pruning, distillation)
- Understanding of target hardware including mobile SoCs and desktop/laptop GPUs
- Strong Python for export pipelines and training-side tooling
- Working fluency with deployed models
- Collaborative working style with clear communication and reliable delivery
Nice to have
- Experience shipping world-model, neural-rendering, or real-time generative pipelines on device
- Hands-on experience deploying models through WebGPU including writing or tuning WGSL compute shaders
- Game-engine or real-time-graphics background (Unity, Unreal, or custom engine)
- Contributions to open-source ML inference frameworks or runtimes
- Familiarity with compiler stacks for custom kernel generation and graph optimization
- Experience with on-device benchmarking infrastructure and performance-regression CI
- Proficiency in C++, Objective-C, or Swift for runtime integration