LogoLanguage
Way Dot Com (P) Ltd

Way.Com India Pvt Ltd, 4th Floor, Yamuna, Technopark Phase-III, Thiruvananthapuram , 695583 , 695583

Site Reliability Lead Engineer

Closing Date:13,Sept 2026
Job Published: 10,Sept 2026
Contact Email: careers@way.com

Brief Description

Apply Now:https://waydot.greythr.com/hire/jobs/site-reliability-lead-engineer

About Us

Way.com — America’s dominant automotive super app, powering every mile for over 10 million customers—from auto insurance and parking to EV charging and more. Leveraging cutting-edge AI, machine learning, and advanced data analytics, we deliver innovative, personalized solutions that transform car ownership. Featured by Bloomberg and ranked 48th on Andreessen Horowitz’s Marketplace list of fastest-growing companies worldwide and recognized as the top Product in Insurance by UnitQ, we’re making car ownership easier, more affordable, and more rewarding.

Role: Site Reliability Lead Engineer 

We're looking for a Site Reliability Lead Engineer to own the reliability, availability, and performance of our production platform while leading a small team of SRE and production support engineers. You'll set the technical direction for observability, incident response, and automation - staying hands-on with the systems while making the team and platform around you more resilient. 

Responsibilities 

  • Lead a team of SRE and production support engineers across incident response, on-call, and reliability engineering 
  • Own SLIs, SLOs, and error budgets for the platform, and drive the engineering work needed to meet them 
  • Design and build automation for deployment, monitoring, alerting, and self-healing across CI/CD pipelines 
  • Lead major incident response and drive blameless post-incident reviews through to root-cause fixes 
  • Own the observability stack (ELK, metrics, tracing) and improve signal quality to reduce mean time to detection and recovery 
  • Optimize database (MySQL/MongoDB), application (Java/Spring Boot, Tomcat), and infrastructure (Docker, Linux) performance 
  • Manage the 24/7 on-call rotation and continuously reduce operational toil and alert fatigue 
  • Partner with engineering teams to build reliability and operability into features before they ship 
  • Mentor engineers on reliability practices, incident command, and production ownership 

Preferred Skills

Required Qualifications 

  • 10+ years of software engineering or SRE experience, including operating large-scale production systems 
  • Deep experience with Linux system administration, service management, log analysis, and performance troubleshooting 
  • Strong backend fundamentals in Java and Spring Boot, plus Tomcat and Docker in production 
  • Hands-on experience with the ELK stack and modern monitoring, alerting, and tracing tooling 
  • Production experience with Kafka (producers/consumers, troubleshooting, and operational support) 
  • Solid database operations across MySQL and MongoDB, including performance optimization 
  • Proven experience owning CI/CD pipelines and deployment automation (Jenkins or equivalent) 
  • Team leadership experience: leading engineers, running on-call rotations, and driving incident response 
  • Strong communication skills with both technical and non-technical stakeholders 

Apply Now:https://waydot.greythr.com/hire/jobs/site-reliability-lead-engineer