Listed on Ando’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
A research role building evaluation frameworks and benchmarks for AI agents in a messaging platform, working with real production data to measure whether agent behavior improvements actually help teams. Best suited for researchers who have shipped evaluation systems and can bridge the gap between offline benchmarks and live product performance.
Our summary, not Ando’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Strong applied research background in model evaluation, benchmarking, or failure analysis
- Work samples or code demonstrating eval frameworks or benchmark suites
- Strong technical communication skills
- Ability to work with messy, incomplete, ambiguous real-world data
Nice to have
- Familiarity with simulation techniques (Park et al.)
- Knowledge of human-in-the-loop evaluation methods
- Experience with context/memory-compression work
- Familiarity with agent observability tools like Langsmith
- Experience with multi-party long-running settings and memory systems