Required
- Should work in US hours ( 6am Pacific Time to 5 Pacific Time )
- 4+ years administering production Hadoop clusters (HDP, CDH, or equivalent).
- Deep, hands-on HDFS, YARN, and Hive administration — not just usage. You should be comfortable reading NameNode and metastore logs and reasoning about JVM heap behavior.
- Strong Linux systems administration: process and memory forensics, SSH, systemd, disk and network troubleshooting on RHEL/CentOS-family hosts.
- Spark-on-YARN operational experience, including memory overhead tuning and diagnosing driver-side failures.
- Solid scripting in Bash and Python.
- Judgment about shared infrastructure. We want someone who will say "that benchmark measured database throughput, not gateway capacity — start lower and ramp" instead of maximizing a number.
- Ambari administration and in-place Hadoop version upgrades.
- Apache Airflow, particularly AWS MWAA.
- MySQL and MongoDB operations at scale.
- Elasticsearch or Sphinx/Manticore search infrastructure.
- Terraform and AWS (S3, EMR, IAM).
- Experience decommissioning or migrating legacy Hadoop estates onto newer platforms.
Send your resume to careers@polussolutions.com
Please include “Hadoop Administrator & Data Engineer - [Your Name]” in the subject line.