Data Engineer at Solve Intelligence

on-site · full_time · Visa sponsorship

Apply for this role at Solve Intelligence

Responsibilities:
- Build and operate high-throughput, resumable ingestion pipelines with incremental updates, monitoring and failure recovery.
- Parse and extract structured content from diverse document types while preserving metadata and handling malformed records and evolving schemas.
- Design schemas, indexes and partitioning for keyword, vector and structured search over tens to hundreds of millions of records.
- Link and reconcile records across patents, scientific literature and other sources; preserve provenance, dates and versions.
- Profile and tune parsing, ingestion, database builds and queries; diagnose CPU, memory and storage I/O bottlenecks and optimise for throughput, latency and cost.
- Collaborate closely with AI researchers and product engineers; take end-to-end ownership from source acquisition to serving queries.

Requirements:
- Strong proficiency in Python and SQL and demonstrable experience operating production databases.
- Hands-on experience building and running production data pipelines for large, messy datasets.
- Practical experience operating search systems (keyword and vector) over large document collections.
- Solid understanding of schema design, indexing, query optimisation and performance profiling.
- Proven track record of diagnosing and fixing performance bottlenecks in live systems through measurement and tuning.
- Nice-to-have: experience with PostgreSQL/pgvector, OpenSearch/Elasticsearch, Spark/Delta Lake, AWS, NoSQL or systems languages (Rust/C++).