LogoLanguage
Quantiphi Analytics Solutions Private Limited

6th floor, Carnival Technopark, Technopark Campus, Karyavattom, Thiruvananthapuram, Kerala 695581 , 695581

Engagement Manager - Support Lead

Closing Date:16,Oct 2026
Job Published: 01,Sept 2026

Brief Description

What you’ll do:  

As an Engagement Manager and Techno-Functional Support Lead for the Production Engineering (PE) and Service Delivery (SD) stream, you will be the primary interface between our critical healthcare AI platforms and our key stakeholders, both internal and external. You will own the strategic direction, reliability, scalability, and availability of these platforms, moving beyond day-to-day ticket resolution to define and implement a comprehensive Support Strategy rooted in Site Reliability Engineering (SRE) principles.

You will lead a high-performing team of engineers, guiding them through complex incident resolutions, orchestrating advanced cloud automation, and expertly bridging the gap between Data Engineering, ML Engineering, Operations, and our business stakeholders. This role demands a unique blend of deep technical acumen, strategic thinking, and exceptional client-facing communication skills. You will serve as the primary escalation point, technical architect, and relationship manager for support operations, ensuring high availability and optimal performance for multi-geography projects while aligning technical solutions with business objectives.

Key Responsibilities:

Strategic Engagement & Leadership:

  • Client & Stakeholder Management: Serve as the primary technical and functional point of contact for key clients and internal business units regarding platform reliability, performance, and service delivery. Translate complex technical issues and resolutions into clear, business-centric communications.

  • Support Strategy & Roadmap: Define, champion, and execute the long-term support strategy, incorporating SRE principles (SLOs, SLIs, Error Budgets) to proactively enhance platform reliability and user experience. Align this strategy with overall product and business goals.

  • Team Leadership & Development: Lead, mentor, and empower a team of Senior and Junior Support Engineers. Foster a culture of continuous improvement, technical excellence, and client-centricity. Conduct performance reviews, manage shift rosters, and facilitate knowledge sharing.

  • Major Incident Management & Communication: Act as the Major Incident Manager and the primary escalation POC. Lead technical resolution war rooms, drive Root Cause Analysis (RCA), and own the Corrective Action processes. Crucially, manage all stakeholder communications during critical incidents, providing timely and transparent updates.

Technical Architecture & Automation:

  • SRE Implementation & Governance: Drive the adoption and continuous improvement of SRE practices, defining and tracking key metrics (SLOs, SLIs, Error Budgets) to ensure platform reliability and performance meet business expectations.

  • Infrastructure as Code (IaC) & Automation: Architect and enforce IaC best practices using Terraform/CloudFormation, ensuring environments are self-healing, reproducible, and compliant. Oversee the design and implementation of comprehensive monitoring, logging, and alerting stacks to enable proactive support.

  • Release & Deployment Management: Oversee the maturity of the release pipeline, ensuring zero-downtime deployments, automated rollbacks, and robust change management processes.

  • Cloud-Native Innovation: Introduce and integrate advanced cloud-native tools and services to improve operational efficiency, enhance automation, and drive cost optimization.

Collaboration & Influence:

  • Design for Supportability: Collaborate closely with Solution Architects, Product Managers, and Engineering teams to influence the design of new features and platforms, ensuring they are inherently supportable, maintainable, and scalable from inception.

  • Cross-Functional Alignment: Bridge the gap between Data Engineering, ML Engineering, and Operations, ensuring seamless integration and efficient resolution of cross-functional issues.

  • Process Optimization: Continuously evaluate and optimize support processes, leveraging automation and best practices to improve efficiency and reduce Mean Time To Resolution (MTTR).

 

Preferred Skills

Required Technical Skills:

Expert Proficiency in Cloud Infrastructure (AWS/GCP preferred):

  • Compute & Serverless: Deep expertise in configuring and optimizing EC2/GCE instances, Lambda/Cloud Functions, and auto-scaling groups based on custom metrics.

  • Networking: Advanced knowledge of VPC peering, Transit Gateways, Load Balancers (ALB/NLB), Route53/Cloud DNS, and VPN/Direct Connect setups.

  • Security & IAM: Ability to design least-privilege IAM policies, manage KMS keys for encryption, and implement WAF rules to protect public-facing endpoints.

  • Storage Lifecycle: Managing tiered storage strategies (S3/GCS lifecycle policies) and high-performance block storage optimization (EBS/Persistent Disks).

Software Engineering Proficiency:

  • Core Development: Strong command of Python (using frameworks like Flask/Django/FastAPI) or Java (Spring Boot) to build production-grade applications and tooling.

  • White-Box Debugging: Ability to clone application repositories, navigate complex codebases, attach remote debuggers, and identify the exact line of code causing logic errors or memory leaks.

  • Code Contribution & Review: Comfortable submitting Pull Requests (PRs) for bug fixes, conducting thorough code reviews for peers, and writing robust unit/integration tests to prevent regressions.

  • Performance Profiling: Experience using profilers (e.g., cProfile, JProfiler) to analyze thread dumps and heap dumps to resolve performance bottlenecks.

Microservices & API Architecture:

  • Protocols & Patterns: Deep understanding of RESTful API design, gRPC protobufs, and asynchronous communication patterns (Pub/Sub, Kafka, SQS).

  • Resiliency Patterns: Implementation of Circuit Breakers, Retry logic (exponential backoff), and Rate Limiting to prevent cascading failures.

  • Service Mesh: Hands-on experience with traffic splitting, mutual TLS (mTLS), and observability sidecars.

  • Distributed Tracing: Analyzing request flows across microservices using tools like OpenTelemetry to pinpoint latency hotspots.

Container Orchestration (Kubernetes - GKE/EKS preferred):

  • Cluster Administration: Managing GKE/EKS upgrades, node pools, and understanding the control plane components (API Server, Scheduler, Controller Manager, etcd).

  • Networking & Security: Configuring Ingress Controllers (Nginx/ALB), Network Policies (Calico/Cilium), and Pod Security Policies/Contexts.

  • Resource Management: Tuning resource requests/limits, setting up Horizontal/Vertical Pod Autoscalers (HPA/VPA), and debugging errors.

  • Helm & Operators: Writing and managing complex Helm charts and understanding how Kubernetes Operators automate stateful applications.

Automation & Tooling:

  • Internal Tools Development: Building custom CLI tools or web dashboards to empower L1/L2 support teams (e.g., a "one-click" user reset tool or a log analyzer dashboard).

  • Event-Driven Automation: Creating self-healing workflows (e.g., a Lambda function that automatically restarts a hung process or cleans up temporary files when disk alerts fire).

Infrastructure as Code (IaC):

  • Terraform Mastery: Structuring Terraform projects using reusable modules, managing remote state locking (S3 + DynamoDB), and using workspaces for multi-environment management.

  • Configuration Management: Using Ansible playbooks for OS-level hardening, patch management, and configuration drift detection.

  • Policy as Code: Implementing tools like OPA (Open Policy Agent) or Sentinel to enforce compliance checks within the IaC pipeline (e.g., preventing public S3 buckets).

Databases (SQL & NoSQL):

  • Relational (PostgreSQL): Analyzing outputs to optimize slow queries, managing connection pooling (PgBouncer), and configuring replication/WAL archiving for disaster recovery.

  • NoSQL (Cassandra/MongoDB): Understanding consistency levels, partition keys, sharding strategies, and diagnosing compaction or garbage collection issues.

  • Caching: Implementing and troubleshooting caching layers (Redis/Memcached) to offload database read pressure.

DevOps Toolchain:

  • CI/CD Pipelines: Designing complex Jenkins pipelines (Groovy shared libraries) that include parallel build stages, automated testing, and canary deployments.

  • Version Control: Advanced Git strategies (Gitflow/Trunk-based), handling large merge conflicts, and using git hooks for pre-commit checks.

  • Artifact Management: Managing Docker registries and dependency repositories (Artifactory/Nexus) with lifecycle policies to clean up old builds.