Machine Learning Engineer, Infra, AI for Drug Discovery at Roche | CA, US | Rezi

Machine Learning Engineer, Infra, AI for Drug Discovery at Roche

Machine Learning Engineer, Infra, AI for Drug Discovery

Roche · CA, US

1 months ago

Machine Learning Engineer, Infra, AI for Drug Discovery

Roche · CA, US

2 months ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Machine Learning Engineer, Infra, AI for Drug Discovery role.

Rezi rewrites your resume against Roche's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Machine Learning Engineer, Infra, AI for Drug Discovery posting at Roche — free, in seconds.

About the Role

The Computational Sciences Center of Excellence (CoE) is a strategic, unified group focused on harnessing the power of data and Artificial Intelligence (AI) to assist scientists in drug discovery and development. The AI for Drug Discovery (AI4DD) group is building machine learning platforms to enable researchers and engineers to move models from experimentation into reliable scientific and production workflows. This role will contribute to the model-serving platform and the broader infrastructure required to make machine learning models easier to deploy, scale, observe, and safely incorporate into scientific and agentic workflows.

Responsibilities

  • Design, implement, ship, and operate scalable model-serving infrastructure for machine learning, scientific, LLM, and agentic workloads.
  • Help evolve the internal model deployment platform into a reliable, self-service platform for teams across the organization.
  • Improve platform scalability and reliability, including scale-to-zero, faster model startup, workload isolation, traffic management, and reduction of request failures and latency bottlenecks.
  • Build observability and operational tooling for model usage, latency, reliability, resource consumption, inference cost, bottlenecks, and service-level indicators.
  • Improve the usability of model deployment by developing validated configuration interfaces, reusable deployment patterns, APIs, command-line tools, and documentation.
  • Help converge real-time and batch inference workflows onto shared platform capabilities where appropriate.
  • Contribute to model lifecycle management infrastructure, including model registration and versioning, evaluation, promotion and release gates, monitoring, environment progression, and rollback.
  • Build event-driven integrations that connect model publication, evaluation, promotion, deployment, and retraining workflows.
  • Build consistent metrics and evaluation signals for understanding model cost, quality, reliability, and fitness for downstream workflows.
  • Partner with machine learning, data, scientific, and platform teams to translate requirements into maintainable solutions and remove infrastructure bottlenecks.
  • Own workstreams from design through implementation and production support, using strong software-engineering practices including testing, reviews, documentation, and incremental delivery.

Requirements

  • BS or MS in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
  • 3+ years of relevant industry experience in software engineering, infrastructure engineering, platform engineering, DevOps, MLOps, or a related area.
  • Strong Python programming skills and experience building and shipping maintainable production software, services, automation, or developer tooling.
  • A demonstrated interest in hands-on implementation and production software delivery.
  • Experience designing, deploying, or operating cloud systems (preferably on AWS) using services such as EKS, EC2, S3, IAM, SQS, SNS, and CloudWatch.
  • Experience with containers, Kubernetes, Helm, and IaC tools such as Terraform or Pulumi.
  • Experience with CI/CD, Git-based development workflows, automated testing, and software release practices.
  • Ability to troubleshoot complex systems using metrics, logs, traces, events, and observability tools such as Datadog, Prometheus, Grafana, or OpenTelemetry.
  • Understanding of distributed-systems concepts such as concurrency, queuing, retries, timeouts, idempotency, backpressure, and failure recovery.
  • Ability to gather requirements, communicate technical tradeoffs, and document systems for users and engineers with varied infrastructure experience.
  • Demonstrated ability to independently deliver practical, incremental solutions while considering immediate needs and longer-term platform direction.

Skills

  • Python
  • AWS
  • EKS
  • EC2
  • S3
  • IAM
  • SQS
  • SNS
  • CloudWatch
  • Kubernetes
  • Helm
  • Terraform
  • Pulumi
  • CI/CD
  • Git
  • Datadog
  • Prometheus
  • Grafana
  • OpenTelemetry
  • KServe
  • Triton
  • vLLM
  • Ray Serve
  • Prefect
  • Dagster

Location

  • California
  • New York

Work Type

  • Onsite

Experience Level

  • 3+ years of relevant industry experience

Education Level

  • BS or MS in Computer Science, Engineering, or a related technical field, or equivalent practical experience.

Salary/Compensations

  • California: $147,600 - $274,000
  • New York: $141,100 - $262,100

Benefits

  • Discretionary annual bonus
  • Benefits detailed at the link provided below

About the Company

  • Roche’s Research and Early Development organisations at Genentech (gRED) and Pharma (pRED) are leveraging AI, data, and computational sciences to accelerate R&D.
  • The new Computational Sciences Center of Excellence (CoE) is a strategic, unified group whose goal is to harness the transformative power of data and Artificial Intelligence (AI) to assist scientists in both pRED and gRED to deliver more innovative and transformative medicines for patients worldwide.
  • At Roche’s AI for Drug Discovery (AI4DD) group (Prescient Design), we are building the machine learning platforms that enable researchers and engineers to move models from experimentation into reliable scientific and production workflows.

Equal Opportunity

  • Genentech is an equal opportunity employer. It is our policy and practice to employ, promote, and otherwise treat any and all employees and applicants on the basis of merit, qualifications, and competence.
  • The company's policy prohibits unlawful discrimination, including but not limited to, discrimination on the basis of Protected Veteran status, individuals with disabilities status, and consistent with all federal, state, or local laws.
  • If you have a disability and need an accommodation in relation to the online application process, please contact us by completing this form Accommodations for Applicants.