Senior Machine Learning Engineer
Babbel
NEW
Seniority
Senior
Model
In-Office
Sector
Salary
Undisclosed
Contract
Full-Time
Babbel's learner-personalisation engine tracks what a learner has and hasn't mastered, and decides what they should practice next. This is a hands-on senior individual-contributor role on that team. You will own real subsystems end to end and make decisions about how to improve the personalisation engine, from research and benchmarking through production, monitoring and incidents.
What you'll do
- Work directly with the Principal Scientist to take designs from spec into production, then keep improving what you've built on your own judgement rather than waiting to be told what's next.
- Help shape new features from the beginning — not just implementation of a spec handed to you.
- Take real ownership of core personalisation and mastery-tracking subsystems, operate independently, and make decisions confidently and competently.
- Design the evaluation that tells you whether a change is real: a rigorous offline benchmark against a real baseline, and the online experiment that confirms or kills it.
- Run what you build. Instrument it, notice when it's silently wrong rather than only when it errors, and fix it before it becomes an incident.
- Deliver with coding agents as a matter of course, and verify what they produce before you rely on it.
What you'll need
- Strong, hands-on ML engineering that has shipped real models to production — recommendation, ranking, scoring, or trust-and-safety systems under real user load are the closest match.
- Experience with probabilistic modeling, latent-variable modeling and Bayesian inference, or the equivalent rigor from an adjacent domain.
- ML system evaluation: monitoring metrics you define, debugging output that doesn't look right, and rolling out a change to a live scoring or ranking system without breaking it.
- Rigorous experimentation practice: benchmarking against a real baseline, running or reading A/B tests correctly, and the judgement to know when an offline improvement won't survive contact with production.
- TypeScript/Python as your primary languages, with enough command of AWS, Terraform, and CI/CD to ship and own your own service's delivery.
- Coding agents are part of your daily workflow, and you check their output before you rely on it.
Nice to have
- Experience with psychometric models, such as Item Response Theory.
- Graph ML experience — embeddings, graph neural networks, or relational modeling — at real scale.
- Public technical work: open-source contributions, writing, or competitive ML.
