Key Responsibilities
Data Pipeline & Lakehouse Engineering
- Build and maintain scalable ETL/ELT pipelines to ingest data from clinical trial management systems and third-party vendors.
- Manage and optimize the Databricks Lakehouse architecture, including the implementation of the Medallion Architecture (Bronze, Silver, and Gold layers).
- Perform expert-level SQL development and performance tuning on SQL Server 2016.
- Transition legacy batch-processing jobs into modern, event-driven, or request-response API patterns to ensure data is accessed, not copied.
API-Centric Integration Development
- Design and deploy robust APIs that allow our Databricks Lakehouse to consume data directly from Medidata, Salesforce, Workday, Certinia, and NetSuite.
- Build API endpoints that enable these enterprise systems to query the Lakehouse for real-time clinical insights and financial metrics.
- Implement Databricks Unity Catalog as the governance layer to manage and secure these cross-system API connections
Customer Data Access & Sharing
- Develop and maintain the API suite used by our customers to access their clinical data.
- Implement Databricks Delta Sharing to provide customers with secure, direct access to live data tables without the need for traditional ETL or file transfers.
- Create comprehensive API documentation and developer portals to facilitate seamless data consumption by customer systems.
Data Modeling & Analytics Enablement
- Develop and maintain data models within the lakehouse to support advanced analytics and clinical reporting.
- Partner with analysts, AI engineers, and data scientists to curate datasets.
Data Governance, Quality & Security
- Assist in the design and implementation of data quality and monitoring solutions.
- Work within Microsoft Purview to map data lineage across API boundaries, ensuring transparency in how data is accessed across the organization.
- Ensure all API integrations meet strict security standards, including OAuth2, OpenID Connect, and industry-specific compliance like GxP and HIPAA.
- Monitor API performance and latency to ensure that data access remains as fast as needed.
- Create and maintain documentation for data lineage, schemas, and pipeline logic.
Required Qualifications
- Bachelor’s degree in Computer Science, Information Systems, or a related technical field.
- Must have at least 3-5 years of professional experience in data engineering, data warehousing, or systems integration with a proven track record of proficiency.
- Strong proficiency in SQL, with specific experience in SQL Server 2016 environments.
- Hands-on experience with Databricks, including Delta Lake, Unity Catalog, and Delta Sharing.
- Demonstrated ability to build ETL pipelines using traditional tools (such as SSIS) or modern methods (such as Azure Data Factory, dbt, and/or Lakeflow).
- Expert-level experience building and consuming RESTful APIs using Python, C#, or Scala.
- Hands-on experience with the API ecosystems of major platforms including Salesforce, Workday, and NetSuite.
- Thorough understanding of data modeling principles, including dimensional modeling and relational design.
- Experience managing data within cloud storage environments like Azure Data Lake.
- Clear understanding of zero-copy architecture and how to use APIs and data virtualization to provide system-to-system visibility without duplicating underlying datasets.
- Familiarity with Microsoft Purview and/or Unity Catalog for metadata management.

