Data Engineer at EMBL
Hinxton, United Kingdom
on-site · full_time · 3303 · Visa sponsorship
Responsibilities:
- Analyse and optimise existing data pipelines for performance, scalability and reliability.
- Design and implement ETL processes and new pipeline architectures for large-scale structural biology datasets.
- Integrate pipelines with PDBe, PDBe-KB, AFDB and internal tools in collaboration with bioinformaticians and annotators.
- Monitor pipeline health, investigate incidents, tune queries and resolve bottlenecks to maintain throughput.
- Lead and execute database migration tasks (notably Oracle → PostgreSQL) ensuring data integrity and minimal disruption.
- Document pipelines, workflows, runbooks and operational procedures for handovers and knowledge sharing.
- Research and recommend tools, techniques and best practices to improve data infrastructure and cost/performance trade-offs.
Requirements:
- MSc in Computer Science, IT, Bioinformatics or related field with demonstrated IT expertise.
- Strong data modelling skills and advanced SQL capability across complex schemas.
- Proficient in Python and experienced with ETL tooling for large-volume data processing.
- Hands-on experience with relational databases: deep PostgreSQL knowledge (architecture, partitioning, indexing, tuning), Oracle (PL/SQL) and MySQL/MariaDB.
- Proven record of migrating databases between RDBMS platforms, especially Oracle to PostgreSQL.
- Familiarity with data warehousing concepts and cloud analytics (Redshift, BigQuery) is desirable.
- Good communicator able to work effectively in interdisciplinary, international teams.