Advisor - Data Architect, Data Foundry at Lilly | CA, US | Rezi

Advisor - Data Architect, Data Foundry at Lilly

Advisor - Data Architect, Data Foundry

Lilly · CA, US

3 weeks ago

Advisor - Data Architect, Data Foundry

Lilly · CA, US

23 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Lilly Small Molecule Discovery is dedicated to creating molecules that improve lives. The Discovery Technology and Platforms (DTP) team accelerates this process by developing optimized platforms, integrating advanced technologies and data connectivity to streamline lab operations, and investing in new capabilities. Data Foundry, a key team within DTP, drives AI-native drug discovery through four integrated pillars: Architecture4Insight, Methods4Insight, Automation & Scale4Insight, and Preparedness4Insight. This role is crucial for building the data infrastructure that enables AI-native drug discovery, transforming raw scientific data into machine-actionable, FAIR-compliant, and insight-ready assets for both human scientists and AI agents.

Responsibilities

  • Design and implement data models, schemas, and ontologies for chemical, biological, and automation-generated data to support discovery workflows.
  • Define and maintain controlled vocabularies, metadata standards, and FAIR-compliant data frameworks.
  • Implement semantic data standards (RDF, OWL, SPARQL) and ontology engineering practices for interoperable, machine-readable scientific data.
  • Design and implement data lakehouse architecture using modern platforms (Databricks, Snowflake, or equivalent), including data storage patterns, partitioning strategies, and query optimization.
  • Build and optimize ETL/ELT pipelines using Spark, dbt, or similar tools.
  • Implement real-time and streaming data integration (Kafka, Kinesis, event-driven patterns) connecting LIMS, instruments, and lab automation systems.
  • Design and implement knowledge graphs (Neo4j, Amazon Neptune, TigerGraph) to capture relationships across the discovery landscape.
  • Architect specialized data solutions such as array databases (TileDB), document stores (MongoDB), and vector databases.
  • Build query and traversal patterns to enable scientists and AI agents to ask relational questions across the data landscape.
  • Partner with scientific software engineers to ensure data architectures are implementable, performant, and well-documented.
  • Collaborate with Methods4Insight to design data structures supporting analytical model training, deployment, and evaluation.
  • Work with Tech@Lilly to define scaling strategies, ensure enterprise compliance, and transition data architectures to production.
  • Contribute to build-versus-buy-versus-adopt decisions by evaluating commercial and open-source data platforms.

Requirements

  • M.S. or PhD in Computer Science, Data Science, Bioinformatics, Computational Biology, Information Science, or related STEM field.
  • MS (with 6+ years) and PhD (with 2+ years) of data architecture, data engineering, or scientific informatics experience.
  • Deep expertise in at least one of the focus areas: relational databases, data modeling and ontology engineering, data platform and lakehouse architecture (Databricks, Snowflake, Spark), or knowledge graph and specialized database systems (Neo4j, Neptune, MongoDB, TileDB).
  • Working familiarity with multiple database paradigms — relational, graph, document, columnar, key-value — and strong SQL skills.
  • Understanding of scientific data types and experimental workflows in life sciences or pharma (chemical, biological, HTE data).
  • Strong communication skills with ability to translate data architecture concepts for both technical and scientific audiences.
  • Familiarity with cloud platforms (AWS, Azure, or GCP) and modern data integration patterns.
  • Pharmaceutical or biotech research industry experience, particularly in discovery data management or research informatics.
  • Experience with semantic web technologies: RDF, OWL, SPARQL, Protégé, or equivalent ontology engineering tools.
  • Hands-on experience with graph databases (Neo4j, Neptune, TigerGraph) and knowledge graph design patterns for scientific data.
  • Data lakehouse architecture experience: Databricks (Delta Lake, Unity Catalog), Snowflake, or equivalent; ETL/ELT with Spark, dbt.
  • Experience with streaming/real-time data platforms (Kafka, Kinesis, Flink) and event-driven architectures.
  • Familiarity with LIMS, ELN systems (e.g., Benchling), and laboratory instrument data integration.
  • Experience with vector databases (Pinecone, Weaviate, pgvector) and embedding-based retrieval for ML/RAG applications.
  • Array database experience (TileDB, Zarr) for genomics, imaging, or high-dimensional scientific data.
  • FAIR data principles implementation experience and Data Readiness Level frameworks.
  • Scientific data standards and controlled vocabularies in chemistry (InChI, SMILES) or biology (Gene Ontology, UniProt).
  • Experience with C, C++, or Rust for performance-critical data processing; familiarity with HPC data I/O patterns for large-scale scientific computations.

Skills

  • Data Modeling
  • Ontology Engineering
  • Data Architecture
  • Data Engineering
  • Scientific Informatics
  • Relational Databases
  • Lakehouse Architecture
  • Knowledge Graphs
  • Specialized Databases
  • SQL
  • Cloud Platforms (AWS, Azure, GCP)
  • Data Integration
  • Semantic Web Technologies (RDF, OWL, SPARQL)
  • Graph Databases (Neo4j, Neptune, TigerGraph)
  • ETL/ELT
  • Spark
  • dbt
  • Streaming Data Platforms (Kafka, Kinesis, Flink)
  • Event-Driven Architectures
  • LIMS
  • ELN Systems
  • Vector Databases
  • Array Databases (TileDB, Zarr)
  • FAIR Data Principles
  • Scientific Data Standards
  • C, C++, Rust

Location

  • San Francisco, CA

Work Type

  • Full-time

Experience Level

  • MS (6+ years)
  • PhD (2+ years)

Education Level

  • M.S. or PhD in Computer Science, Data Science, Bioinformatics, Computational Biology, Information Science, or related STEM field

Salary/Compensations

  • $151,500 - $244,200

Benefits

  • Company bonus (depending, in part, on company and individual performance)
  • Company-sponsored 401(k)
  • Pension
  • Vacation benefits
  • Medical, dental, vision and prescription drug benefits
  • Flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts)
  • Life insurance and death benefits
  • Certain time off and leave of absence benefits
  • Well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities)

About the Company

  • At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters.
  • Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve.
  • This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.

Equal Opportunity

  • Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions.
  • Lilly is proud to be an EEO Employer and does not discriminate on the basis of age, race, color, religion, gender identity, sex, gender expression, sexual orientation, genetic information, ancestry, national origin, protected veteran status, disability, or any other legally protected status.
  • Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees. Our current groups include: Africa, Middle East, Central Asia (AMECA), Black Employees at Lilly (BE@Lilly), Chinese Culture Network (CCN), EnAble, Evolve, Lilly Indian Network (LIN), Organization of Latinx at Lilly (OLA), Pride (LGBTQ+ Allies), Veterans Leadership Network (VLN) and Women’s Initiative for Leading at Lilly (WILL).