About the Role
Drive the design, deployment, and operationalisation of machine learning solutions on Google Cloud, bridging AI/ML engineering and cloud delivery to ensure reliable, secure, and scalable production models and pipelines.
Responsibilities
- Design and implement end-to-end MLOps pipelines on Vertex AI, including data ingestion, model training, evaluation, and deployment.
- Build and manage Vertex AI Pipelines (Kubeflow Pipelines) for automated model training and retraining workflows.
- Deploy and manage models using Vertex AI Model Registry, Endpoints, and Batch Prediction services.
- Implement feature engineering workflows using Vertex AI Feature Store.
- Develop GCP-native integrations connecting Vertex AI with BigQuery, Dataflow, Cloud Storage, and Pub/Sub.
- Manage infrastructure for ML workloads using Terraform, ensuring reproducible and version-controlled environments.
- Configure IAM policies for Vertex AI workloads including service account governance and VPC Service Controls.
- Lead cloud delivery activities: sprint planning, release management, environment promotion, and stakeholder communication.
- Establish model monitoring using Vertex AI Model Monitoring for data drift and skew detection.
- Collaborate with data scientists to containerise experiments and promote models through dev/staging/production.
- Drive adoption of GKE for model serving workloads where custom inference infrastructure is required.
Requirements
- 4+ years of experience with GCP, including 2+ years hands-on with Vertex AI.
- Strong proficiency in Python and ML frameworks (TensorFlow, PyTorch, Scikit-learn).
- Experience building Vertex AI Pipelines and managing model lifecycle in Vertex AI Model Registry.
- Solid Terraform skills for provisioning Vertex AI, GCS, BigQuery, and associated infrastructure.
- Good understanding of GCP IAM, particularly for securing ML pipelines and data access.
- Experience with GCP-native development patterns and event-driven architectures.
- Demonstrated cloud delivery experience including planning, execution, and stakeholder management.
- Familiarity with containerisation (Docker) and GKE for model serving.
- Google Professional Machine Learning Engineer certification.
- Experience with LLM fine-tuning, Vertex AI Generative AI Studio, or Model Garden.
- Familiarity with Feast, Tecton, or similar feature stores.
- Experience with Ansible for environment configuration and automation.
- Background in DataOps or platform engineering for data-intensive workloads.
Skills
- Vertex AI
- MLOps
- Kubeflow Pipelines
- Vertex AI Model Registry
- Vertex AI Endpoints
- Vertex AI Batch Prediction
- Vertex AI Feature Store
- BigQuery
- Dataflow
- Cloud Storage
- Pub/Sub
- Terraform
- GCP IAM
- VPC Service Controls
- Docker
- GKE
- Python
- TensorFlow
- PyTorch
- Scikit-learn
- LLM fine-tuning
- Vertex AI Generative AI Studio
- Model Garden
- Feast
- Tecton
- Ansible
- DataOps
- Platform Engineering
Experience Level
- 4+ years of experience with GCP
- 2+ years hands-on with Vertex AI
