- Clanx
- Bengaluru, KA
- Full-Time
- 41 days ago
Senior Applied Machine Learning Engineer (Eval) - Bengaluru.
Before you go
Before you leave us, sign up for our email alerts
We don't do job spam, just the best digital jobs delivered straight to your inbox.
Senior Applied Machine Learning Engineer (Eval) - Bengaluru: our view in 3 lines...
- The Role:This role is for a senior applied machine learning engineer focused on AI quality for an agentic video editor.
- The Person:The person will define quality standards, build evaluation datasets and offline and online evaluations, analyse agent failures, and improve agent quality through better data, model selection, fine-tuning, and regression checks.
- Requirements:The ideal candidate has strong software engineering skills in TypeScript or Python, experience shipping an LLM or agent system used by real customers, and experience building evaluations, datasets, experiments, or AI quality systems.
About the role
Senior Applied ML Engineer to own AI quality for Cardboard’s agentic video editor, building evaluation datasets, offline/online evals, regression checks, and feedback loops that turn production failures into measurable improvements.
Company Details
Cardboard is an AI-first video editor building agentic tools that understand user requests, work with media, and make real edits on the timeline. It is backed by a Tier-1 global fund, YC, and founders of billion-dollar companies. Website: https://www.usecardboard.com
Requirements
-
Experience shipping and operating an LLM or agent system used by real customers.
-
Strong software engineering skills in TypeScript or Python, with ability to work across both.
-
Experience building evaluations, datasets, experiments, or AI quality systems.
-
Strong product judgment and ability to turn vague AI quality issues into measurable problems.
-
Ability to work across data, evaluation methods, model selection, and fine-tuning.
-
Strong ownership as a senior individual contributor.
-
Bonus: Experience with multimodal AI, video, media, or creative software.
-
Bonus: Experience with human labeling, model graders, or fine-tuning.
-
Bonus: Strong understanding of experiment design and statistics.
Responsibilities
-
Define quality standards for Cardboard’s agent.
-
Build trusted evaluation datasets from real product usage.
-
Build offline and online evaluations using automated checks, model graders, and human review.
-
Analyze real agent runs and identify recurring failure patterns.
-
Improve agent quality through better data, evaluation methods, model selection, and fine-tuning.
-
Build regression checks and release gates for important agent changes.
-
Track AI quality alongside latency and cost.
-
Partner with product and engineering teams to ship measurable improvements.
Job Details
Bengaluru, India
Interview Process
-
Recruiter Screen
-
Technical Interview
-
ML & Evaluation Deep Dive
-
Product & Engineering Interview
-
Final Interview
Important Note
ClanX is a recruitment partner, helping Cardboard hire Senior Applied ML Engineer, Evals & Data.

