Impress employers and recruiters.
Choose from hundreds of resume examples.

Impress employers and recruiters.
Choose from hundreds of resume examples.
Tailor your resume to this Software Engineer, Model Inference, DeepMind role.
Rezi rewrites your resume against Google's job description. Free.

Tailor your resume to this Software Engineer, Model Inference, DeepMind role.
Rezi rewrites your resume against Google's job description. Free.
Don't guess if your resume is good enough.
See how it scores against the Software Engineer, Model Inference, DeepMind posting at Google — free, in seconds.

Don't guess if your resume is good enough.
See how it scores against the Software Engineer, Model Inference, DeepMind posting at Google — free, in seconds.
About the Role
At Google DeepMind, our mission is to build the world's first general-purpose learning agent. In this role, you will be at the forefront of bringing AI research to life, optimizing and deploying large language models (LLMs) like Gemini onto Google's production infrastructure. This involves a blend of technical expertise and collaborative problem-solving to ensure efficiency and quality throughout the LLM deployment lifecycle. The role offers opportunities for both individual contributor and team lead roles, and is open to Software Engineering and Research Engineering backgrounds.
Responsibilities
- Collaborate closely with Research teams to understand next generation modeling approaches, ensuring they are designed and implemented with production considerations in mind.
- Work with infrastructure teams to deliver serving infrastructure that is designed for maximum efficiency and performance, addressing bottlenecks in speed, scale, and quality.
- Identify opportunities to automate tasks, eliminate redundancies, build performant tests, and improve the overall velocity of model releases.
- Gain a deep understanding of serving frameworks, pre-processing pipelines, caching mechanisms, and other relevant technologies.
- Leverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers (XLA), custom kernels (Pallas), and serving infrastructure on hardware accelerators (TPUs/GPUs).
Requirements
- Bachelor’s degree or equivalent practical experience.
- 8 years of experience in software development.
- 2 years of experience in deploying and maintaining machine learning models in a live production environment.
- Experience in profiling, configuring, or executing ML workloads directly on hardware accelerators (e.g., GPU or TPU).
- Experience designing, building, or optimizing model serving infrastructure or inference backends.
Skills
- Developing serving infrastructure
- Programming hardware accelerators (GPUs, TPUs) via ML frameworks (e.g., JAX, PyTorch) or low-level programming models (e.g., Pallas, CUDA, OpenCL)
- Profiling software to identify performance bottlenecks
- Distributed ML systems optimization and parallelism (e.g., data, model, or pipeline parallelism)
- Writing performance-optimized kernels
- Understanding of LLM architecture and inference performance dynamics (e.g., Transformer models, memory bandwidth and compute bounds, KV cache scaling)
Location
- USA
Work Type
- Full-time
Experience Level
- 8 years of experience in software development
- 2 years of experience in deploying and maintaining machine learning models
Education Level
- Bachelor’s degree or equivalent practical experience
About the Company
- At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users.
- We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.
- We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.