Staff Site Reliability Engineer
Doctolib
Seniority
Senior
Model
In-Office
Sector
Salary
Undisclosed
Contract
Full-Time
Staff Site Reliability Engineer joining the SRE team dedicated to platform reliability within Platform Engineering. Your mission will be to act as a technical leader driving Doctolib's reliability and scalability at a European scale, ensuring our platform remains reliable, debuggable, and resilient across infrastructure, observability, and cross-cutting reliability initiatives.
What you'll do
- Lead large-scale cross-cutting reliability initiatives across the platform, spanning infrastructure automation, observability, and incident management
- Identify and drive improvements to incident detection, response, and postmortem analysis capabilities
- Define and evolve SLOs, error budgets, and alerting standards across multiple product teams
- Take part in the on-call rotation, and actively contribute to improving our on-call experience by refining alerting, reducing noise, and ensuring actionable telemetry
- Serve as a mentor and technical coach to senior engineers, helping elevate the craft of reliability engineering across the company
- Influence strategic decisions by providing technical guidance to leadership and representing reliability engineering in architectural reviews and platform discussions
- Partner with software engineering teams to embed reliability practices early in the development lifecycle
What you'll need
- Extensive experience (8+ years) in SRE, platform engineering, or infrastructure roles within a large-scale, multi-team production environment
- Proven experience with cloud platforms such as AWS, GCP, or Azure
- Strong experience with containerization and orchestration technologies, Kubernetes is a must, its deployment and scaling strategies ecosystem
- Implemented and operated SLIs, SLOs, and error budgets in production
- Experience managing on-call rotations and leading incident response in high-stakes environments
- Strong systems engineering background with fluency in at least one backend programming language (e.g., Go, Python, Ruby)
- Proven ability to lead through influence: setting technical direction, driving consensus, and mentoring engineers across teams
- Fluent in English
Nice to have
- Deep expertise in observability tooling and architecture (logging, tracing, metrics)
- Experience designing and operating high-scale telemetry pipelines and working with developers to improve instrumentation quality
- Experience in regulated environments, healthcare, fintech, or similar
- Care about reliability enablement, golden paths, runbooks, shared libraries
What they offer
- Deutschlandticket (Germany-wide public transport pass) fully paid
- 28 vacation days + 1 additional day for each full calendar year of employment (up to 30 days)
- Company health insurance with supplementary benefits through Allianz
- Company pension scheme (bAV) with 40% employer subsidy
- Free mental health and coaching services through Moka.care
- Flexible workplace policy offering hybrid and office-based modes

