What you’ll do:
As an Engagement Manager and Techno-Functional Support Lead for the Production Engineering (PE) and Service Delivery (SD) stream, you will be the primary interface between our critical healthcare AI platforms and our key stakeholders, both internal and external. You will own the strategic direction, reliability, scalability, and availability of these platforms, moving beyond day-to-day ticket resolution to define and implement a comprehensive Support Strategy rooted in Site Reliability Engineering (SRE) principles.
You will lead a high-performing team of engineers, guiding them through complex incident resolutions, orchestrating advanced cloud automation, and expertly bridging the gap between Data Engineering, ML Engineering, Operations, and our business stakeholders. This role demands a unique blend of deep technical acumen, strategic thinking, and exceptional client-facing communication skills. You will serve as the primary escalation point, technical architect, and relationship manager for support operations, ensuring high availability and optimal performance for multi-geography projects while aligning technical solutions with business objectives.
Key Responsibilities:
Strategic Engagement & Leadership:
-
Client & Stakeholder Management: Serve as the primary technical and functional point of contact for key clients and internal business units regarding platform reliability, performance, and service delivery. Translate complex technical issues and resolutions into clear, business-centric communications.
-
Support Strategy & Roadmap: Define, champion, and execute the long-term support strategy, incorporating SRE principles (SLOs, SLIs, Error Budgets) to proactively enhance platform reliability and user experience. Align this strategy with overall product and business goals.
-
Team Leadership & Development: Lead, mentor, and empower a team of Senior and Junior Support Engineers. Foster a culture of continuous improvement, technical excellence, and client-centricity. Conduct performance reviews, manage shift rosters, and facilitate knowledge sharing.
-
Major Incident Management & Communication: Act as the Major Incident Manager and the primary escalation POC. Lead technical resolution war rooms, drive Root Cause Analysis (RCA), and own the Corrective Action processes. Crucially, manage all stakeholder communications during critical incidents, providing timely and transparent updates.
Technical Architecture & Automation:
-
SRE Implementation & Governance: Drive the adoption and continuous improvement of SRE practices, defining and tracking key metrics (SLOs, SLIs, Error Budgets) to ensure platform reliability and performance meet business expectations.
-
Infrastructure as Code (IaC) & Automation: Architect and enforce IaC best practices using Terraform/CloudFormation, ensuring environments are self-healing, reproducible, and compliant. Oversee the design and implementation of comprehensive monitoring, logging, and alerting stacks to enable proactive support.
-
Release & Deployment Management: Oversee the maturity of the release pipeline, ensuring zero-downtime deployments, automated rollbacks, and robust change management processes.
-
Cloud-Native Innovation: Introduce and integrate advanced cloud-native tools and services to improve operational efficiency, enhance automation, and drive cost optimization.
Collaboration & Influence:
-
Design for Supportability: Collaborate closely with Solution Architects, Product Managers, and Engineering teams to influence the design of new features and platforms, ensuring they are inherently supportable, maintainable, and scalable from inception.
-
Cross-Functional Alignment: Bridge the gap between Data Engineering, ML Engineering, and Operations, ensuring seamless integration and efficient resolution of cross-functional issues.
-
Process Optimization: Continuously evaluate and optimize support processes, leveraging automation and best practices to improve efficiency and reduce Mean Time To Resolution (MTTR).

