Job Description
We re looking for a Site Reliability Lead Engineer to own the reliability, availability, and performance of our production platform while leading a small team of SRE and production support engineers. You ll set the technical direction for observability, incident response, and automation - staying hands-on with the systems while making the team and platform around you more resilient.
Responsibilities
Lead a team of SRE and production support engineers across incident response, on-call, and reliability engineering
Own SLIs, SLOs, and error budgets for the platform, and drive the engineering work needed to meet them
Design and build automation for deployment, monitoring, alerting, and self-healing across CI/CD pipelines
Lead major incident response and drive blameless post-incident reviews through to root-cause fixes
Own the observability stack (ELK, metrics, tracing) and improve signal quality to reduce mean time to detection and recovery
Optimize database (MySQL/MongoDB), application (Java/Spring Boot, Tomcat), and infrastructure (Docker, Linux) performance
Manage the 24/7 on-call rotation and continuously reduce operational toil and alert fatigue
Partner with engineering teams to build reliability and operability into features before they ship
Mentor engineers on reliability practices, incident command, and production ownership