Bioinformatics Data Engineer (RNA Resources) at EMBL

Hinxton, United Kingdom

on-site · full_time · 3303 · Visa sponsorship

Apply for this role at EMBL

Responsibilities:
- Run, maintain and optimise data pipelines for Rfam and RNAcentral to ensure reliable, scalable processing and fast retrieval.
- Modernise and containerise curation pipelines (Docker/Singularity) and implement reproducible, human-in-the-loop AI-assisted curation workflows.
- Design, develop and scale LLM-driven pipelines for literature summarisation, curation and agentic workflows.
- Build scalable workflows for ncRNA annotation in genomes and support data release cycles and validation steps.
- Produce clear documentation, tests and reproducible deployment artifacts; collaborate with developers, curators and project leads.
- Represent the resources at conferences and consortium/SAB meetings; gather and act on community feedback.

Requirements:
- Master’s degree (or equivalent) in computational biology, bioinformatics or a related discipline.
- Strong Python skills and experience with bioinformatics tool development; familiarity with Bash/Unix; Rust or other languages desirable.
- Proven experience developing and operating production bioinformatics pipelines using Nextflow, Snakemake or similar.
- Solid SQL expertise and experience with PostgreSQL/MySQL including performance tuning, partitioning, indexing and query optimisation.
- Practical experience with containerisation (Docker/Singularity), HPC environments and workflow reproducibility.
- Experience building or integrating LLMs/AI technologies (e.g., LangChain or comparable tooling) into production workflows.
- Good communication skills and willingness to engage with scientific users and present technical work.