About the Role
We are seeking a Senior DevOps Engineer to join our Applied AI practice, focusing on the intersection of platform engineering and AI delivery. This role involves leading the optimization and evolution of cloud infrastructure, deployment pipelines, and operational practices to deliver high-quality AI systems for clients in enterprise environments.
Responsibilities
- Architect, build, and continuously enhance CI/CD pipelines to automate and accelerate software delivery.
- Lead the management and optimisation of cloud infrastructure (AWS), ensuring scalability, security, and reliability.
- Design, implement, and maintain Infrastructure as Code (IaC) with tools such as Terraform and CloudFormation.
- Proactively monitor, troubleshoot, and enhance system performance, availability, and security.
- Drive the adoption of containerisation and orchestration technologies like Docker and Kubernetes.
- Improve system observability by implementing advanced logging, monitoring, and alerting.
- Lead the implementation of security best practices, including IAM, secrets management, and vulnerability assessments.
- Collaborate closely with developers to continuously optimise build, deployment, and scaling strategies.
- Automate key operational tasks and apply SRE principles to enhance system reliability.
- Take ownership of incident response and lead root cause analysis for production issues.
- Implement prompt versioning, model evaluation pipelines, and controlled promotion gates for LLMOps.
- Design observability for token costs, inference latency, retrieval quality, and model drift detection.
- Build agentic resilience by implementing rate limiting, circuit breakers, and graceful fallbacks for LLM APIs.
- Own inference cost engineering, including throughput management, caching strategy, and cost-per-query alerting.
- Design AI-native CI/CD pipelines with evaluation harnesses and golden dataset regression tests.
Requirements
- 5+ years of hands-on experience in DevOps, SRE, or Cloud Engineering.
- Extensive expertise in AWS cloud platforms and services.
- Practical experience with Kubernetes and containerisation technologies.
- Strong scripting and automation skills with Bash, Python, or Go.
- In-depth knowledge of CI/CD tools including Jenkins, GitHub Actions, GitLab CI/CD, and ArgoCD.
- Solid experience with Infrastructure as Code tools including Terraform and CloudFormation.
- Comprehensive understanding of Linux administration and networking fundamentals.
- Experience implementing security best practices including IAM, SSL/TLS, and compliance frameworks such as SOC2, ISO 27001, and GDPR.
- Proficiency in monitoring and logging tools including the ELK Stack, Prometheus, Grafana, or Datadog.
- Exceptional problem-solving skills and the ability to operate in a fast-moving, ambiguous environment.
- Strong communication and collaboration skills to work effectively across cross-functional teams, including client stakeholders.
- Unrestricted working rights in Australia.
Skills
- AWS
- Kubernetes
- Docker
- Terraform
- CloudFormation
- Bash
- Python
- Go
- Jenkins
- GitHub Actions
- GitLab CI/CD
- ArgoCD
- Linux Administration
- Networking
- IAM
- SSL/TLS
- SOC2
- ISO 27001
- GDPR
- ELK Stack
- Prometheus
- Grafana
- Datadog
- LLMOps
- Serverless architectures (AWS Lambda)
- Database performance tuning
- AI/ML workloads
- LangSmith
- Weave
Location
- Sydney
Work Type
- Hybrid
- Full-time
Experience Level
- Senior
Benefits
- Monthly anchor days
- Team lunches
- Annual offsite
- Unlimited access to Go1's learning library
- Support from internal performance coach
About the Company
- Bilue is a digital consultancy that designs and builds smart, user-friendly technology for well-known Australian businesses.
- Our culture is people-first and purpose-driven, with a down-to-earth, values-led team.
- We value excellence, not ego, and foster a high-trust environment with low politics.
- We have offices in Sydney and Melbourne, and a growing presence in Manila.
