Our client is a technological company at the forefront of the global medical research scene with real-world data. As a part of their growth, they need a Senior Data Engineer whose main task would be the development of a Data Platform to unify the company’s data flow.
Responsibilities:
> Designing and managing data engineering workflows.
> Developing and maintaining ETL pipelines to process and analyze large-scale datasets.
> Implementing and optimizing data storage in data warehousing systems.
> Ensuring data accuracy and accessibility.
Requirements (Must):
> Extensive experience building custom data collectors and parsers with Python.
> Hands-on experience with Apache Airflow to understand how to manage task dependencies at scale and monitor complex DAGs.
> Experience with implementing and maintaining Airbyte (OSS) and dbt for scalable ETL processes.
> Deep expertise in both Row-based Relational DBs (for application/transactional data) and Columnar/ OLAP DBs (for high-performance analytics).
> Strong proficiency with Docker and Linux environments.
> Extreme ownership. You don’t just wait for tickets; you identify bottlenecks, push for solutions, and ensure data integrity from the initial source to the final ML model.
Brownie points:
> Previous experience with SAP HANA or ClickHouse is a major plus.
> Experience building feature stores or specialized pipelines designed specifically for ML consumption (MLOps).
> Experience using Terraform or Ansible.
> Familiarity with healthcare data standards (e.g., HL7, FHIR) or experience working within HIPAA/GDPR-compliant environments.