Skip to main content
CareerApp

Skill

Model Evaluation

Data Science, Analytics and AI/ML

The practice of systematically assessing a machine learning or AI model's performance using metrics like accuracy, precision, recall, F1 score, or task-specific benchmarks. Data scientists and ML engineers use it to compare candidate models, detect overfitting, and validate that a model meets quality and fairness standards before deployment. It typically involves held-out test sets, cross-validation, and increasingly, human or automated evaluation of qualitative outputs for generative models.

See who is hiring

Model Evaluation in the job market

Last checked September 12, 2026

Open roles
18

on Career App right now

Employers
14

hiring for it

Median pay

not enough disclosed

Disclose pay

of these roles

Open roles requiring Model Evaluation (18)

Bristol Myers Squibb

Senior AI Engineer

Bristol Myers Squibb

Full-time · $137,530 – $183,319

This is a hands-on senior AI engineering role building cloud-native, AI-powered applications and agentic systems for pharmaceutical use cases. The position suits experienced backend engineers who are comfortable with frontier AI technologies, want to work on high-impact problems in a fast-paced delivery model, and can bridge AI capabilities with production reliability and enterprise architecture.

Listed on Bristol Myers Squibb’s careers site · Apply there ↗

Staff, Data Science & Applied AI

Warner Bros. Discovery

Atlanta, GA

You'll design and deploy production machine learning models and generative AI solutions across Warner Bros. Discovery's streaming, gaming, and content platforms, working at the intersection of statistical rigor and business impact. This senior individual contributor role suits experienced data scientists who thrive building scalable analytics systems and translating complex business problems into measurable enterprise value.

Listed on Warner Bros. Discovery’s careers site · Apply there ↗

Vice President, Product - AI Center of Excellence

Mastercard

Full-time · New York, NY · $245,000 – $391,000

This Vice President role leads product strategy for an enterprise AI platform that enables internal teams to safely build, deploy, and scale AI agents. It suits experienced product leaders who have built platforms in regulated industries and can bridge technical AI capabilities with governance, risk management, and business value.

Listed on Mastercard’s careers site · Apply there ↗

Product Manager, AI

LSEG

New York, NY · $126,900 – $211,500

This role blends product management with hands-on AI engineering, building agentic systems that operate on financial data at scale. It suits someone with deep technical roots in ML or software who has shipped LLM-powered products and can own both the "what" and the "how" of complex AI systems.

Listed on LSEG’s careers site · Apply there ↗

Forward Deployed Machine Learning Engineer

Protege

Remote

A machine learning engineer role focused on building evaluation infrastructure and benchmarks for AI models at an early-stage data platform company. This suits engineers with ML experience who want to own full-stack infrastructure problems and work closely with customers and researchers to shape a new product vertical.

Listed on Protege’s careers site · Apply there ↗

Senior Staff Software Engineer, Perception (R4985)

Shield AI

Full-time · Washington, DC · $233,760 – $350,640

This senior engineer role involves building and deploying machine learning models that help autonomous defense systems perceive and understand their environment, bridging research and production. It suits experienced ML engineers who want to work on vision and vision-language models at scale while solving real-world deployment challenges for military applications.

Listed on Shield AI’s careers site · Apply there ↗

Research Engineer

Ando

Full-time · San Francisco, CA

A research role building evaluation frameworks and benchmarks for AI agents in a messaging platform, working with real production data to measure whether agent behavior improvements actually help teams. Best suited for researchers who have shipped evaluation systems and can bridge the gap between offline benchmarks and live product performance.

Listed on Ando’s careers site · Apply there ↗

AI Engineer

MovieMagic

Tempe, AZ

Entertainment Partners is hiring a senior AI engineer to develop and deploy machine learning models and AI systems for its entertainment technology products. This role suits someone with strong PyTorch and transformer expertise who can move models from research through production while maintaining software engineering rigor.

Listed on MovieMagic’s careers site · Apply there ↗

Apply for Principal Machine Learning Engineer, Data Mining at Motional on their site

Principal Machine Learning Engineer, Data Mining

Motional

Full-time · Boston, MA · $144,000 – $192,000 · Remote

This role builds machine learning systems that automatically discover critical edge cases and failures in autonomous vehicle sensor data, helping improve safety through smarter data mining and analysis. It suits engineers who want to apply foundation models and representation learning to solve real-world problems at scale in a rapidly advancing industry.

Listed on Motional’s careers site · Apply there ↗

Product Manager, Claude Code Model Performance

Menlo Ventures Portfolio

Full-time · San Francisco, CA · $305,000 – $460,000

This role leads the performance and launch strategy for Claude Code, Anthropic's AI coding agent, by building evaluations, coordinating with research and engineering teams, and translating model improvements into capabilities developers actually need. It suits PMs with hands-on experience building AI evaluations, deep familiarity with coding models, and the ability to bridge research breakthroughs with shipped products.

Listed on Menlo Ventures Portfolio’s careers site · Apply there ↗

View all 18 on Jobs

Asked for alongside Model Evaluation

Measured from the 18 open roles that name Model Evaluation — not from a curated list.

Employers hiring for Model Evaluation

14 in total, most open roles first.

Related skills

Curated neighbors in the taxonomy, whether or not employers ask for them together.

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.