Senior Data Engineer at Formation Bio | Boston, MA, US | Rezi

Senior Data Engineer at Formation Bio

Senior Data Engineer

Formation Bio · Boston, MA, US

5 days ago

Senior Data Engineer

Formation Bio · Boston, MA, US

5 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Senior Data Engineer role.

Rezi rewrites your resume against Formation Bio's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Senior Data Engineer posting at Formation Bio — free, in seconds.

About the Role

As a Senior Data Engineer at Formation Bio, you will build trusted data systems supporting clinical operations, drug asset evaluation, business development, analytics, and machine learning. You will work across various data sources to design and operate reliable ingestion pipelines, transformations, data models, and data products, shaping how application data is modeled and exposed.

Responsibilities

  • Design and operate production data systems that ingest clinical, operational, and third-party vendor data into reliable, queryable data products.
  • Own shared and canonical data models, data contracts, transformations, orchestration, warehouse models, and downstream interfaces.
  • Partner with Product Engineering on application data models, source-system contracts, APIs, events, and data access patterns.
  • Partner with Data Science on productionized training datasets, feature pipelines, data interfaces, and ML use cases.
  • Turn recurring data cleaning, normalization, and transformation work into versioned, tested, observable, and maintainable production pipelines.
  • Build data products for clinical operations, asset evaluation, Business Development, analytics, machine learning, and AI Enabled Employees and Agents.
  • Design data products that are semantically clear, discoverable, machine-readable, permission-aware, traceable, and safe to query.
  • Establish strong data quality, testing, freshness, completeness, lineage, documentation, and observability practices.
  • Own data governance practices for sensitive and regulated data, including access controls, auditability, traceability, and appropriate data handling.
  • Participate in support and incident response for data platform issues, including diagnosis, stakeholder communication, remediation, and prevention of recurrence.
  • Use AI tools, including LLMs and agentic coding systems, to accelerate pipeline development, data modeling, debugging, documentation, and data quality investigation while validating their output.
  • Contribute to architecture and design reviews, mentor other engineers, and improve the engineering practices used across the organization.

Requirements

  • 5+ years of relevant data engineering experience building and operating production data systems.
  • Experience with pharmaceutical, biology, HIPPA or other regulated data core to BioTech is required.
  • Strong Python and SQL skills, with deep experience in data modeling and warehouse systems, especially Snowflake.
  • Experience with Dagster as an orchestration system or equivalent transformation tooling such as dbt.
  • Experience with data contracts, schema evolution, data quality testing, observability, lineage, and production incident response.
  • Experience integrating messy clinical, operational, vendor, or otherwise complex source data.
  • Working knowledge of Docker, GitHub, and Terraform or OpenTofu sufficient to partner effectively with SRE.
  • Experience building data products and access patterns for applications, Data Science, analytics, human users, and AI Enabled Employees and Agents.
  • Strong judgment about when to build reusable platform capabilities versus one-off stakeholder solutions.
  • Daily fluency with AI tools and the ability to validate generated code, transformations, and data-modeling decisions.
  • Exceptional collaboration and communication skills across Product Engineering, Data Science, Clinical Operations, Data Management, Business Development, and other non-technical partners.
  • Experience working within and building validated computerized systems (CSV) is a plus.

Skills

  • Python
  • SQL
  • Snowflake
  • Dagster
  • dbt
  • Docker
  • GitHub
  • Terraform
  • OpenTofu
  • AI tools
  • LLMs
  • Agentic coding systems

Location

  • New York City metro area
  • Boston metro area
  • Research Triangle (NC)
  • San Francisco Bay Area

Work Type

  • Hybrid

Experience Level

  • Senior

Salary/Compensations

  • $185,500 - $232,000

Benefits

  • Equity
  • Comprehensive benefits
  • Generous perks

About the Company

  • Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development.
  • Advancements in AI and drug discovery are creating more candidate drugs than the industry can progress because of the high cost and time of clinical trials.
  • Formation Bio, founded in 2016 as TrialSpark Inc., has built technology platforms, processes, and capabilities to accelerate all aspects of drug development and clinical trials.
  • The company partners, acquires, or in-licenses drugs from pharma companies, research organizations, and biotechs to develop programs past clinical proof of concept and beyond, ultimately helping to bring new medicines to patients.
  • The company is backed by investors across pharma and tech, including a16z, Sequoia, Sanofi, Thrive Capital, John Doerr, Spark Capital, SV Angel Growth, and others.
  • At Formation Bio, our values are the driving force behind our mission to revolutionize the pharma industry.
  • Every team and individual at the company shares these same values, and every team and individual plays a key part in our mission to bring new treatments to patients faster and more efficiently.

Equal Opportunity

  • Formation Bio is committed to building a diverse and inclusive team.
  • We are an equal opportunity employer and welcome candidates from all backgrounds.
  • All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, national origin, ancestry, sex (including pregnancy, childbirth, breastfeeding, and related medical conditions), gender identity or expression, sexual orientation, age, disability, genetic information, marital status, military or veteran status, or any other characteristic protected by federal, state, or local law.