About the Role
Isomorphic Labs is applying frontier AI to help unlock deeper scientific insights, faster breakthroughs, and life-changing medicines with an ambition to solve all disease. The future is coming, enabled and enriched by machine learning, where diseases are curtailed or cured starting with better and faster drug discovery. Join an interdisciplinary team driving groundbreaking innovation and play a meaningful role in achieving ambitious goals within an inspiring and collaborative culture.
Responsibilities
- Develop and operate large-scale bioinformatics pipelines for high-throughput data analysis, ensuring reliable processing from raw data to ML-ready datasets.
- Apply bioinformatics best practices to the ingestion and harmonization of complex datasets, ensuring model training is grounded in high-quality, version-controlled biological data, and coherently integrated datasets.
- Harmonise disparate public databases (e.g., Ensembl, UniProt, Reactome, Open Targets), implementing rigorous versioning and mapping strategies to mitigate identifier collisions, data loss, and semantic drift across releases.
- Act as a strategic partner to ML Research, Computational Biology, Drug Development, and Chemistry teams, championing the adoption of the internal bioinformatics platform and standardized biological data primitives into their daily research workflows.
- Participate in research projects as a "Deployed Engineer", providing customized solutions, and identifying technical gaps, while ensuring project-specific insights are contributed back into the core bioinformatics platform.
- Provide documentation, guidance, and training on data resources and curation processes to the wider organization.
Requirements
- Proven experience with the large-scale processing of raw bioinformatics data (e.g., FASTQ, BAM, mzXML).
- A demonstrable track record of delivering high-quality bioinformatics outputs across varied modalities (e.g., genomics, proteomics, functional genomics, systems biology, single cell).
- Experience delivering bioinformatics solutions directly to research teams, scientific communities, or industry projects, with a strong focus on user enablement.
- Experience writing production-grade code in Python and developing automated, scalable bioinformatics pipelines.
- PhD or MSc in Bioinformatics, Computational Biology, or a related field, or equivalent practical experience in a biopharmaceutical or research environment.
- Experience with domain-specific workflow systems (e.g., Nextflow) for scaling high-throughput pipeline execution.
- Familiarity with general-purpose data orchestration and processing frameworks (e.g., Dagster, Apache Beam) for integrating research pipelines into a production platform.
- Familiarity with building and maintaining bioinformatics infrastructure on Google Cloud Platform (GCP).
- Familiarity with modern, high-performance DataFrame libraries (e.g., Polars), and relational data modeling and analysis (SQL).
- Exposure to machine learning concepts and the specific data requirements for training ML models.
- Demonstrable experience working with regulated PHI data.
- Extensive experience in software development with Python.
Skills
- Python
- Bioinformatics
- Computational Biology
- Data Harmonization
- Data Curation
- Machine Learning
- Genomics
- Proteomics
- Functional Genomics
- Systems Biology
- Single Cell Analysis
- Nextflow
- Dagster
- Apache Beam
- Google Cloud Platform (GCP)
- Polars
- SQL
- Software Development
Location
- Hybrid
Work Type
- Hybrid
Experience Level
- PhD or MSc in Bioinformatics, Computational Biology, or a related field, or equivalent practical experience in a biopharmaceutical or research environment.
- Proven experience with large-scale processing of raw bioinformatics data.
- Demonstrable track record of delivering high-quality bioinformatics outputs.
- Experience delivering bioinformatics solutions directly to research teams, scientific communities, or industry projects.
- Experience writing production-grade code in Python.
- Extensive experience in software development with Python.
Education Level
- PhD or MSc in Bioinformatics, Computational Biology, or a related field, or equivalent practical experience in a biopharmaceutical or research environment.
About the Company
- Isomorphic Labs (IsoLabs) was launched in 2021 to advance human health by building on and beyond the Nobel-winning AlphaFold system.
- Our interdisciplinary team of drug discovery experts and machine learning specialists has built powerful new predictive and generative AI models that accelerate scientific discovery at digital speed.
- Our name comes from the belief that there is an underlying symmetry between biology and information science.
- By harnessing AI’s powerful capabilities, we can use it to model complex biological phenomena to help design novel molecules, anticipate how drugs will perform and develop innovative medicines to treat and cure some of the world’s most devastating diseases.
- We have built a world-leading drug design engine comprising AI models that are capable of working across multiple therapeutic areas and drug modalities.
- We are continually innovating on model architecture and developing cutting-edge capabilities to advance rational drug design.
- Every day, and with each new breakthrough, we’re getting closer to the promise of digital biology, and achieving our ambitious mission to one day solve all disease with the help of AI.
Equal Opportunity
- We are committed to equal employment opportunities regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, pregnancy or related condition (including breastfeeding) or any other basis protected by applicable law.
- If you have a disability or additional need that requires accommodation, please do not hesitate to let us know.
