Location
London
Hours
Full Time
Salary
Competitive salary
About the Role
We’re looking for a Data Engineer with a strong background in bioinformatics and genetics to help build the data pipelines supporting Our Future Health’s growing Clinical Research Recruitment Service. This is an opportunity to combine genetics, bioinformatics and modern data engineering, working with data at significant scale while contributing directly to research that could improve how diseases are prevented, detected and treated. Our Future Health is an ambitious collaboration between the public, charity and private sectors, designed to help people live healthier lives for longer through better prevention, earlier detection and improved treatment of diseases. We will speed up the discovery of new methods of early disease detection, and the evaluation of new diagnostic tools, to help identify and treat diseases early, when outcomes are usually better. With over 2.7M volunteers across the UK, we’re now the world’s biggest health research programme of its kind, and our volunteer group is also more diverse than other, similar health research programmes. Technology and data are central to our mission. Our systems power web sites, clinics across the UK, secure analytics and research systems, pipelines that process highly sensitive health and genetic data, and we are continuing to grow our engineering capability to support this ambition. Our Clinical Research Recruitment Service helps life sciences and academic organisations identify potential participants for clinical studies. As the service grows, we need to turn scientific workflows and manual processes into robust, reusable and scalable production pipelines. We’re looking for a Data Engineer with a solid understanding and experience of bioinformatics, in particular tools and methods associated with genomic data. You can design, build and test pipelines using a range of different technologies. You know how to create repeatable and reusable products and can communicate to and between technical and non-technical stakeholders, with the ability to facilitate discussions and manage different perspectives within a multidisciplinary team including scientists, software engineers, product managers and other data engineers.
Essential Duties and Responsibilities
Support the build of re-usable data pipelines used to identify prospective clinical trial participants. Produce logic for data transformation steps as code, which meets the requirements for our end users and builds well curated, accessible and quality controlled data for analysis. Develop prototypes for pipelines for complex transformations drawing on existing workflows developed in industry and academia. Keep abreast of best practice in data engineering across industry, research and Government and facilitate the adoption of standards. Provide technical input into the upstream parts of the data pipeline, including the specification and transfer of data from data providers. Perform routine ad-hoc data curation activities requiring hands on development of bespoke ETL cleaning scripts using languages such as Python. Work with researchers to understand the data requirements and collaborate to deliver the data needed for their projects.
Experience
- Building and maintaining robust, scalable and efficient data pipelines capable of processing very large amounts of data from multiple systems using a range of technologies.
- Detailed knowledge and understanding of genomic data, with experience in genotyping and imputation advantageous.
- Experience using bioinformatics file standards (VCF, BGEN etc) and tools (PLINK, bcftools, QCtools etc).
- Highly proficient in Python.
- Highly proficient in version control and Git/GitHub.
- Experience with workflow management tools such as Nextflow, WDL/Cromwell, Airflow, Prefect or Dagster.
- Understanding of containerisation technologies (e.g. Docker) and deployment (e.g. Kubernetes).
- Good understanding of cloud environments (ideally Azure), distributed computing and scaling workflows and pipelines.
- Understanding of common data transformation and storage formats, e.g. Apache Parquet.
- Awareness of data standards such as GA4GH and FAIR.
- Experience with Spark, Databricks, and data lakes.
- Follow best practices including code reviews, clean code and unit tests.
- Experience working in an agile development team.
About you
You are pragmatic, collaborative and comfortable working through ambiguity. You enjoy solving difficult problems rather than simply maintaining established systems. You can listen to the needs of technical and business stakeholders, interpret them effectively and manage stakeholder expectations. You are able to communicate clearly to both technical and non-technical audiences and facilitate discussions within multidisciplinary teams.
Qualifications
We welcome applications from all who may not meet every criterion but have most of the experience and skills listed. A strong background in bioinformatics and data engineering is essential. A positive attitude and willingness to learn are highly valued.
Our Future Health




