LogoLanguage
LEADER IT (P) Ltd

SBC 1, -2 FLOOR, THEJASWINI BUILDING, PHASE 1, TECHNOPARK CAMPUS, KARYAVATTOM P O, TRIVANDRUM, KERALA , 695581

Data Engineer (Remote)

Closing Date:18,Sept 2026
Job Published: 09,Sept 2026

Brief Description

Job Title: Data Engineer

Experience: 2–3 Years
Job Type: Full-time (Remote)
Job Function: Data Engineering / Data & AI

Application Process

Interested candidates are required to apply only through the official application link below:

Apply Now: https://leaderit.in/careers/HR-OPN-2026-0032/apply

Important: Applications submitted through any other channel will not be considered. Candidates must apply through the above link to be considered for this position.

Job Overview

We are looking for a skilled and motivated Data Engineer with 2–3 years of relevant experience to join our team. The ideal candidate will have strong hands-on experience in SQL, Python, PySpark, ETL/ELT pipelines, databases, and cloud-based data platforms.

The Data Engineer will work closely with Data Scientists, AI Engineers, Software Developers, Analysts, and Product Teams to build reliable, scalable, secure, and high-quality data solutions. The role involves collecting data from multiple sources, transforming and validating data, building data pipelines, and making data available for analytics, reporting, AI, and machine learning applications.

Key Responsibilities

  • Design, develop, and maintain scalable data pipelines for structured and unstructured data.

  • Build and manage ETL/ELT workflows to collect, clean, transform, validate, and load data.

  • Integrate data from databases, APIs, files, cloud storage, enterprise applications, and third-party platforms.

  • Develop efficient SQL queries, stored procedures, views, and data transformation scripts.

  • Work with relational and NoSQL databases to efficiently store, process, and retrieve data.

  • Design and maintain data models, data warehouses, data lakes, and analytical datasets.

  • Automate data ingestion, transformation, validation, and reporting workflows.

  • Monitor data pipelines and troubleshoot failures, performance issues, and data inconsistencies.

  • Implement data-quality checks to ensure accuracy, completeness, consistency, and reliability.

  • Optimize database queries and data pipelines for performance, scalability, and cost efficiency.

  • Develop reusable data-processing components and maintain technical documentation.

  • Prepare datasets for dashboards, analytics, AI, and machine learning applications.

  • Collaborate with Data Scientists and AI Engineers to provide clean, reliable, and model-ready datasets.

  • Work with Developers and Product Teams to integrate data services into applications and business workflows.

  • Follow data security, privacy, governance, access-control, and compliance requirements.

  • Participate in code reviews, testing, deployment, and production support.

  • Maintain version-controlled code using Git and follow software-development best practices.

Data Pipeline & Platform Responsibilities

  • Build batch and, where required, near-real-time data-processing pipelines.

  • Schedule, monitor, and manage workflows using Apache Airflow or similar orchestration tools.

  • Process large datasets using Python, SQL, Apache Spark, and PySpark.

  • Work with cloud storage, data warehouses, data lakes, and cloud-based data-processing services.

  • Create data models for reporting, dashboards, operational applications, and AI use cases.

  • Implement logging, monitoring, alerting, and error-handling mechanisms for production pipelines.

  • Identify and resolve data-quality issues across source, transformation, and consumption layers.

  • Assist in data migration between legacy systems, cloud platforms, databases, and modern data architectures.

Required Skills

  • Strong proficiency in SQL, including joins, aggregations, subqueries, window functions, and query optimization.

  • Good programming skills in Python for data processing and automation.

  • Hands-on experience developing and maintaining ETL/ELT pipelines.

  • Strong understanding of relational databases such as PostgreSQL, MySQL, SQL Server, or Oracle.

  • Familiarity with NoSQL databases such as MongoDB, Elasticsearch, or similar technologies.

  • Understanding of data modelling, database design, normalization, and dimensional modelling.

  • Knowledge of data formats such as CSV, JSON, XML, Parquet, and Avro.

  • Familiarity with REST APIs and third-party data integrations.

  • Working knowledge of AWS, Microsoft Azure, or Google Cloud Platform.

  • Familiarity with cloud data warehouses, data lakes, and data-processing services.

  • Basic understanding of Apache Spark, Kafka, Airflow, dbt, or similar technologies.

  • Strong knowledge and hands-on experience with PySpark.

  • Experience with Git and collaborative software-development workflows.

  • Understanding of data quality, security, governance, and access-control principles.

  • Strong analytical, troubleshooting, and problem-solving skills.

Preferred Skills

  • Experience with Snowflake, Amazon Redshift, Google BigQuery, Azure Synapse, or Databricks.

  • Familiarity with Docker and containerization.

  • Exposure to CI/CD pipelines and DevOps practices.

  • Experience with real-time data processing and messaging systems such as Apache Kafka.

  • Understanding of Airflow, dbt, or similar data orchestration and transformation frameworks.

  • Familiarity with Power BI, Tableau, Looker, or similar reporting tools.

  • Basic understanding of machine learning and AI data requirements.

  • Experience working with large-scale datasets and enterprise applications.

AI & Emerging Technologies

We are looking for candidates who are interested in supporting modern AI, Machine Learning, and Generative AI applications through reliable data engineering.

The candidate should be willing to:

  • Build and maintain data pipelines for AI and machine learning use cases.

  • Prepare clean, structured, traceable datasets for model training, evaluation, and inference.

  • Work with structured and unstructured data, including text, documents, audio, images, and logs.

  • Support data preparation for vector databases, embeddings, Retrieval-Augmented Generation (RAG), and knowledge-base applications.

  • Explore AI-assisted tools for coding, documentation, testing, debugging, and data analysis.

  • Understand data privacy, security, governance, and responsible AI practices.

  • Continuously learn and adopt emerging data engineering, cloud, AI, and automation technologies.

Familiarity with AI-powered development and productivity tools such as GitHub Copilot, ChatGPT, Gemini, or similar tools will be an advantage.

Qualifications

  • B.Tech, B.E., M.Tech, M.E., MCA, M.Sc. in Computer Science, Information Technology, Data Science, Software Engineering, or a related discipline.

  • 2–3 years of relevant professional experience in Data Engineering, ETL Development, Data Integration, Database Development, or a related field.

  • Hands-on experience building data pipelines and working with databases, APIs, and data-processing technologies.

  • Relevant cloud, database, or data-engineering certifications will be an advantage but are not mandatory.

What We're Looking For

  • Strong foundation in SQL, Python, PySpark, databases, and data pipelines.

  • Ability to understand business requirements and translate them into reliable data solutions.

  • Strong analytical thinking and problem-solving skills.

  • Attention to data accuracy, performance, security, and maintainability.

  • Ability to troubleshoot pipeline failures and data-quality issues.

  • Interest in cloud platforms, AI, analytics, and emerging technologies.

  • Willingness to learn and adapt to new tools and frameworks.

  • Ability to work independently and collaborate effectively with cross-functional teams.

  • Responsible and quality-focused approach to production data systems.

Why Join Us?

  • Work on real-world data, analytics, AI, and enterprise application projects.

  • Gain exposure to modern cloud-based data platforms and engineering practices.

  • Build data pipelines supporting AI, machine learning, dashboards, and business applications.

  • Collaborate with Data Scientists, AI Engineers, Developers, Product Teams, and domain experts.

  • Work with structured and unstructured enterprise data.

  • Learn and work with modern data engineering and AI technologies.

  • Supportive environment for strengthening technical expertise and professional growth.

  • Growth opportunities for candidates demonstrating technical capability, ownership, and initiative.

 

Preferred Skills

Preferred Skills

  • Experience with Snowflake, Amazon Redshift, Google BigQuery, Azure Synapse, or Databricks.

  • Familiarity with Docker and containerization.

  • Exposure to CI/CD pipelines and DevOps practices.

  • Experience with real-time data processing and messaging systems such as Apache Kafka.

  • Understanding of Airflow, dbt, or similar data orchestration and transformation frameworks.

  • Familiarity with Power BI, Tableau, Looker, or similar reporting tools.

  • Basic understanding of machine learning and AI data requirements.

  • Experience working with large-scale datasets and enterprise applications.