Engineering Manager - Site Reliability & Observability
Doctolib
Seniority
Midweight
Model
In-Office
Sector
Salary
Undisclosed
Contract
Full-Time
We are looking for an Engineering Manager to lead an SRE team dedicated to platform reliability and observability within Platform Engineering. Your mission will be to grow and lead a team of Site Reliability Engineers, ensuring Doctolib's platform remains reliable, scalable, and resilient at a European scale across infrastructure, observability, and cross-cutting reliability initiatives.
What you'll do
- Lead, coach, and grow a team of Site Reliability Engineers, supporting their technical development and career progression
- Create a culture of operational excellence, continuous improvement, and psychological safety within the team
- Define and evolve the reliability and observability strategy across the platform, covering infrastructure automation, logging, metrics, tracing, and alerting
- Drive prioritization and roadmap planning for large-scale reliability initiatives, including SLOs, error budgets, and alerting standards across multiple product teams
- Own the team's on-call experience and contribute to incident response processes, ensuring sustainable practices and continuous improvement
- Work closely with Product Managers, Engineering Managers, and architects to align reliability and observability capabilities with product and platform needs
What you'll need
- At least 5+ years of software engineering or SRE experience, with a strong technical background in cloud-native environments (AWS, GCP, and/or Kubernetes-based)
- 3+ years of engineering management experience, leading technical teams (ideally SRE, platform, or infrastructure teams)
- Strong understanding of observability tooling and architecture (e.g. OpenTelemetry, Prometheus, Datadog, Elasticsearch)
- Experience with infrastructure as code (Terraform) and secrets management systems
- Proven ability to balance technical depth with people leadership, able to mentor engineers, review technical designs, and guide architectural decisions
- Fluent in English
Nice to have
- Experience scaling SRE or platform teams in fast-growing, high-traffic environments
- Hands-on experience with backend programming languages (e.g. Go, Python, Ruby)
- Experience driving cultural and technical transformations
- Experience working in regulated environments (healthcare, fintech, or similar)
What they offer
- 28 vacation days + 1 additional day for each full calendar year of employment (up to 30 days)
- Company health insurance with supplementary benefits through Allianz
- Company pension scheme (bAV) with 40% employer subsidy
- ParentCare Program with full salary coverage (100%) first month, 75% second month of birth leave
- Free mental health and coaching services through Moka.care
- Deutschlandticket fully paid, flexible workplace policy, subsidized meals and sports membership

