Skill
Model Evaluation
Data Science, Analytics and AI/ML
The practice of systematically assessing a machine learning or AI model's performance using metrics like accuracy, precision, recall, F1 score, or task-specific benchmarks. Data scientists and ML engineers use it to compare candidate models, detect overfitting, and validate that a model meets quality and fairness standards before deployment. It typically involves held-out test sets, cross-validation, and increasingly, human or automated evaluation of qualitative outputs for generative models.
Model Evaluation in the job market
Last checked September 12, 2026
- Open roles
- 18
- Employers
- 14
- Median pay
- —
- Disclose pay
- —
on Career App right now
hiring for it
not enough disclosed
of these roles
Open roles requiring Model Evaluation (18)
Senior AI Engineer
Bristol Myers Squibb
Full-time · $137,530 – $183,319
This is a hands-on senior AI engineering role building cloud-native, AI-powered applications and agentic systems for pharmaceutical use cases. The position suits experienced backend engineers who are comfortable with frontier AI technologies, want to work on high-impact problems in a fast-paced delivery model, and can bridge AI capabilities with production reliability and enterprise architecture.
Listed on Bristol Myers Squibb’s careers site · Apply there ↗
Staff, Data Science & Applied AI
Warner Bros. Discovery
Atlanta, GA
You'll design and deploy production machine learning models and generative AI solutions across Warner Bros. Discovery's streaming, gaming, and content platforms, working at the intersection of statistical rigor and business impact. This senior individual contributor role suits experienced data scientists who thrive building scalable analytics systems and translating complex business problems into measurable enterprise value.
Listed on Warner Bros. Discovery’s careers site · Apply there ↗
Vice President, Product - AI Center of Excellence
Mastercard
Full-time · New York, NY · $245,000 – $391,000
This Vice President role leads product strategy for an enterprise AI platform that enables internal teams to safely build, deploy, and scale AI agents. It suits experienced product leaders who have built platforms in regulated industries and can bridge technical AI capabilities with governance, risk management, and business value.
Listed on Mastercard’s careers site · Apply there ↗
Product Manager, AI
LSEG
New York, NY · $126,900 – $211,500
This role blends product management with hands-on AI engineering, building agentic systems that operate on financial data at scale. It suits someone with deep technical roots in ML or software who has shipped LLM-powered products and can own both the "what" and the "how" of complex AI systems.
Listed on LSEG’s careers site · Apply there ↗
Forward Deployed Machine Learning Engineer
Protege
Remote
A machine learning engineer role focused on building evaluation infrastructure and benchmarks for AI models at an early-stage data platform company. This suits engineers with ML experience who want to own full-stack infrastructure problems and work closely with customers and researchers to shape a new product vertical.
Listed on Protege’s careers site · Apply there ↗
Senior Staff Software Engineer, Perception (R4985)
Shield AI
Full-time · Washington, DC · $233,760 – $350,640
This senior engineer role involves building and deploying machine learning models that help autonomous defense systems perceive and understand their environment, bridging research and production. It suits experienced ML engineers who want to work on vision and vision-language models at scale while solving real-world deployment challenges for military applications.
Listed on Shield AI’s careers site · Apply there ↗
Research Engineer
Ando
Full-time · San Francisco, CA
A research role building evaluation frameworks and benchmarks for AI agents in a messaging platform, working with real production data to measure whether agent behavior improvements actually help teams. Best suited for researchers who have shipped evaluation systems and can bridge the gap between offline benchmarks and live product performance.
Listed on Ando’s careers site · Apply there ↗
AI Engineer
MovieMagic
Tempe, AZ
Entertainment Partners is hiring a senior AI engineer to develop and deploy machine learning models and AI systems for its entertainment technology products. This role suits someone with strong PyTorch and transformer expertise who can move models from research through production while maintaining software engineering rigor.
Listed on MovieMagic’s careers site · Apply there ↗
Principal Machine Learning Engineer, Data Mining
Motional
Full-time · Boston, MA · $144,000 – $192,000 · Remote
This role builds machine learning systems that automatically discover critical edge cases and failures in autonomous vehicle sensor data, helping improve safety through smarter data mining and analysis. It suits engineers who want to apply foundation models and representation learning to solve real-world problems at scale in a rapidly advancing industry.
Listed on Motional’s careers site · Apply there ↗
Product Manager, Claude Code Model Performance
Menlo Ventures Portfolio
Full-time · San Francisco, CA · $305,000 – $460,000
This role leads the performance and launch strategy for Claude Code, Anthropic's AI coding agent, by building evaluations, coordinating with research and engineering teams, and translating model improvements into capabilities developers actually need. It suits PMs with hands-on experience building AI evaluations, deep familiarity with coding models, and the ability to bridge research breakthroughs with shipped products.
Listed on Menlo Ventures Portfolio’s careers site · Apply there ↗
Asked for alongside Model Evaluation
Measured from the 18 open roles that name Model Evaluation — not from a curated list.
- Machine Learning9 roles50.0%
- Python9 roles50.0%
- Generative AI4 roles22.2%
- LangChain4 roles22.2%
- CI/CD3 roles16.7%
- Data Pipelines3 roles16.7%
- LlamaIndex3 roles16.7%
- Model Deployment3 roles16.7%
Roles that use Model Evaluation
Employers hiring for Model Evaluation
14 in total, most open roles first.
- MPMenlo Ventures Portfolio3 roles
- PPendo2 roles
- UUdacity2 roles
- AAndo1 role
- Bristol Myers Squibb1 role
- FFront1 role
Related skills
Curated neighbors in the taxonomy, whether or not employers ask for them together.