# Research Engineer

Hiring organization: [Ando](https://career.thegoodapps.co/organizations/ando)

Canonical page: https://career.thegoodapps.co/jobs/369cb91d-e708-4024-a616-3778ed6196d0

Listed on Ando's own careers site. Applications go to them directly.

- Employment type: full time
- Location: San Francisco, CA

## Summary

A research role building evaluation frameworks and benchmarks for AI agents in a messaging platform, working with real production data to measure whether agent behavior improvements actually help teams. Best suited for researchers who have shipped evaluation systems and can bridge the gap between offline benchmarks and live product performance.

_Our summary, not Ando's wording._

## Skills named

Benchmarking, Failure Analysis, Model Evaluation

## Required

- Strong applied research background in model evaluation, benchmarking, or failure analysis
- Work samples or code demonstrating eval frameworks or benchmark suites
- Strong technical communication skills
- Ability to work with messy, incomplete, ambiguous real-world data

## Nice to have

- Familiarity with simulation techniques (Park et al.)
- Knowledge of human-in-the-loop evaluation methods
- Experience with context/memory-compression work
- Familiarity with agent observability tools like Langsmith
- Experience with multi-party long-running settings and memory systems

Apply on Ando's site: https://jobs.ashbyhq.com/ando/2d7c68ad-a05a-4906-983e-58699bc08f58/application
