LogoLanguage
Flex and Oncology Clinical Research (P) Ltd

Second Floor, Nila Building, Technopark Phase 1,Kazhakkoottam, Thiruvananthapuram, Kerala , 695581

Data Product Engineer

Closing Date:30,Aug 2026
Job Published: 22,July 2026

Brief Description

Data Product Engineer
Location: Trivandrum - On-site
Department: Data Operations and Technology
 
Role Overview
Worldwide Clinical Trials is seeking a Data Product Engineer to lead the technical evolution of our data assets into high-value, AI-ready products. In this role, you will merge the operational rigor of DataOps with a product-centric mindset, ensuring that our data is discoverable, highly governed, and optimized for Generative AI and Agentic AI workflows. You will be responsible for the production readiness of our data, including governance and master data management, ensuring that data products supporting data intelligence, analytics, Generative AI, and Agentic AI are governed, reliable, and scalable, while also owning the end-to-end lifecycle from ingestion through master data management to secure consumption via Unity Catalog.
 
Key Responsibilities
Data Governance & Master Data Management (MDM)
  • Lead the technical implementation of Master Data Management (MDM) solutions to establish the "Golden Record" for clinical trials, ensuring consistency across disparate data sources.
  • Enforce Data Governance policies through automated "policy-as-code" within the data product lifecycle, ensuring compliance with global clinical regulations.
  • Utilize Databricks Unity Catalog to provide centralized access control, rigorous auditing, and end-to-end data lineage across the entire Azure tenant.
  • Collaborate with data stewards to automate the identification and remediation of sensitive data (PII/PHI) within the data lakehouse.
Data Product Engineering & AI Readiness
  • Treat data as a formal product by defining and maintaining SLAs, SLOs, and SLIs for data quality and availability across the clinical enterprise.
  • Engineer API-based integrations that allow data products to be seamlessly consumed by internal applications, LLMs, and external partners.
  • Ensuring high-quality metadata and are including in data products and available for Generative AI applications.
DataOps Mastery (Data Kitchen & IBM Best Practices)
  • Architect the Data Kitchen "Innovation Pipeline" to allow for the rapid, automated deployment of new data features without compromising production stability.
  • Implement IBM DataOps best practices by creating a unified data virtualization and governance layer that drastically reduces data silos.
  • Automate the "Value Pipeline" using CI/CD for data to ensure that data products are continuously monitored for quality and statistical drift.
  • Foster a culture of "observability" where automated testing is a required part of the data product (governance, MDM) pipeline.
 
Technical Requirements
  • Governance & MDM: Proven experience building MDM frameworks and implementing automated data governance tools in a regulated environment.
  • Cloud Infrastructure: Advanced expertise in the Azure ecosystem, including Azure Data Factory, Azure DevOps, and Azure Key Vault.
  • Databricks Stack: Mastery of Databricks, specifically Unity Catalog for governance and Delta Lake for resilient data lakehouse architecture.
  • Modern Integration: Strong proficiency in PySpark, SQL, and building/consuming RESTful APIs for data exchange programmatic workflow scheduling using Apache Airflow, Databricks Lakeflow, or Prefect.
 
Why Worldwide Clinical Trials?
You will be at the forefront of our data strategy, helping to revolutionize how clinical data is harnessed to bring life-saving treatments to market. We value automation, transparency, and a commitment to quality as we build the "Trusted Factory" for the clinical trials of tomorrow.

Preferred Skills

  • Governance & MDM: Proven experience building MDM frameworks and implementing automated data governance tools in a regulated environment.
  • Cloud Infrastructure: Advanced expertise in the Azure ecosystem, including Azure Data Factory, Azure DevOps, and Azure Key Vault.
  • Databricks Stack: Mastery of Databricks, specifically Unity Catalog for governance and Delta Lake for resilient data lakehouse architecture.
  • Modern Integration: Strong proficiency in PySpark, SQL, and building/consuming RESTful APIs for data exchange programmatic workflow scheduling using Apache Airflow, Databricks Lakeflow, or Prefect.