On-Premise Infrastructure & Self-Managed Kubernetes Focus
This is a full-time Senior Platform & DevOps Engineer position within the Technology & Infrastructure department. The primary focus of the role includes self-managed Kubernetes, on-premise
DMZ/data centers, stateful systems, and CI/CD operations.
Position Overview,
We are seeking an experienced and hands-on Senior Platform & DevOps
Engineer to own, architect, and manage our production infrastructure across
on-premise data centers and hybrid environments. Unlike cloud-native managed
environments EKS/GKE/AKS, this role requires deep bare-metal and virtualized
system expertise, specifically in building, operating, and hardening self-managed
Kubernetes clusters on Linux RHEL/Ubuntu).
You will be responsible for end-to-end platform stability, zero-downtime
deployment pipelines, stateful database cluster operations PostgreSQL,
MongoDB, API Gateway management Kong, high-availability messaging
queues, and robust enterprise network security.
Key Responsibilities,
1. On-Premise Kubernetes & Infrastructure Ownership
● Design, bootstrap, upgrade, and manage bare-metal/VM self-managed
Kubernetes control planes and worker nodes.
● Troubleshoot complex cluster operations, container runtime engines, etcd
backups/restores, CNI networking, ingress controllers, and storage
interfaces CSI/MinIO.
● Enforce node lifecycle management, kernel tuning, security patching, and
OS hardening on RedHat Enterprise Linux RHEL and Ubuntu nodes.
2. Infrastructure as Code & Configuration Management
● Author reusable Terraform and Ansible modules for automated,
reproducible environment provisioning across dev, UAT, production, and
disaster recovery DR environments.
● Eliminate configuration drift and enforce security hardening baselines
across compute, storage, and networking layers.
3. Enterprise DMZ, Security & Network Operations
● Operate within customer-managed DMZ environments, managing
integrations across Web Application Firewalls WAF, load balancers, and
internal/external firewall rules.
● Secure API access using Kong Gateway, managing JWT/HMAC
authentication, TLS/mTLS termination, rate limiting, and RBAC policies.
● Restrict internal workload traffic using Kubernetes NetworkPolicies, private
subnets, and HashiCorp Vault / Kubernetes Secrets management.
4. Stateful Systems & High Availability Operations
● Support stateful infrastructure running on-premise, including
high-availability PostgreSQL clusters, MongoDB Replica Sets, Redis caches,
and RabbitMQ message broker pools.
● Implement and test automated backup, snapshot, and recovery procedures
to guarantee tight Recovery Time Objectives RTO and Recovery Point
Objectives RPO.
5. CI/CD Automation & Observability
● Build and maintain robust GitOps and CI/CD pipelines Jenkins, GitHub
Actions, ArgoCD for microservices packaging, vulnerability scanning
Trivy/SonarQube), and zero-downtime deployment.
● Maintain a central observability stack using Prometheus, Grafana,
Loki/VictoriaLogs, and Alertmanager to proactively monitor golden signals
and alert on system degradation.

