About the Role
The Infrastructure Reliability Engineering Lead is responsible for leading and developing the Infrastructure Reliability Engineering team, ensuring the successful implementation of reliability engineering practices across the infrastructure platforms that underpin LME services. This role translates strategy, standards, and governance into practical engineering outcomes and measurable improvements in service reliability, resilience, and operational risk reduction. The role is accountable for driving engineering excellence across observability, resilience engineering, automation, operational readiness, and service reliability, ensuring infrastructure services are designed, operated, and continuously improved in a secure, recoverable, observable, and supportable manner. The initial emphasis will be on establishing and maturing Infrastructure Reliability Engineering practices, operational readiness assessments, resilience validation, platform observability standards, and reliability-focused automation across critical infrastructure services.
Responsibilities
- Lead, coach, and develop the Infrastructure Reliability Engineering team, setting clear direction, priorities, objectives, and performance expectations.
- Build and maintain a high-performing engineering team through effective recruitment, performance management, coaching, mentoring, and succession planning.
- Allocate engineering resources effectively across reliability improvement initiatives, resilience programs, operational readiness activities, and platform engineering priorities.
- Foster a culture of engineering excellence, accountability, continuous improvement, and shared ownership of service reliability.
- Lead implementation and continual improvement of Infrastructure Reliability Engineering practices across infrastructure and platform services.
- Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), availability metrics, and reliability measures across critical infrastructure services.
- Identify, assess, and drive remediation of reliability risks and systemic weaknesses before they impact production services.
- Lead reliability reviews, trend analysis, and engineering improvement initiatives focused on increasing resilience and reducing operational risk.
- Establish and govern observability standards across infrastructure platforms including monitoring, telemetry, logging, tracing, and alerting.
- Ensure meaningful operational visibility and actionable service health monitoring across infrastructure services.
- Lead improvements in monitoring quality, alert effectiveness, and operational insights through data-driven analysis.
- Ensure infrastructure teams have the observability capabilities required to support effective operational management and rapid issue diagnosis.
- Lead resilience validation activities including failover testing, disaster recovery testing, recovery assurance exercises, and scenario-based resilience assessments.
- Establish and maintain Operational Readiness standards, ensuring services meet agreed supportability, recoverability, observability, and operational acceptance criteria before entering production.
- Ensure recovery capabilities are regularly tested and aligned to agreed business recovery objectives.
- Drive engineering improvements that enhance service continuity, fault tolerance, and recovery capability across critical infrastructure services.
- Lead adoption of Infrastructure as Code, configuration management, and automation practices across infrastructure services.
- Drive standardization, repeatability, and reduction of operational toil through automation and engineering improvements.
- Support the continual improvement of platform reliability, scalability, recoverability, and operational efficiency through engineering-led initiatives.
- Promote best practice engineering approaches across cloud-native, virtualised, and enterprise infrastructure platforms.
- Support the Infrastructure Reliability Engineering Senior Manager in implementing and maturing the Infrastructure Reliability Engineering operating model, standards, and governance frameworks.
- Provide reliability engineering expertise and technical leadership across projects, platform initiatives, and service improvement programs.
- Work collaboratively with Information Security, Service Operations, and Delivery teams to embed reliability, resilience, and operational readiness into platform and service design.
- Manage stakeholder relationships and provide reporting, insights, and recommendations regarding service reliability, resilience, and operational risk.
Requirements
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related technical discipline.
- Significant experience in Infrastructure Reliability Engineering, Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering disciplines.
- Demonstrable experience leading technical engineering teams within complex enterprise environments.
- Candidates should have significant experience operating within complex enterprise technology environments, ideally within financial services or other highly regulated organisations, together with a strong track record of leading teams responsible for critical infrastructure services and engineering outcomes.
- The successful candidate will demonstrate strong people leadership, sound technical judgement, and the ability to influence engineering decisions across multiple technology domains.
- They will possess broad expertise in reliability engineering, resilience, observability, automation, and platform engineering, while remaining capable of contributing hands-on where required.
- Candidates should possess extensive experience improving infrastructure reliability through engineering practices rather than operational intervention, including the use of automation, observability, resilience testing, and service measurement.
- The successful candidate must be capable of providing technical leadership across specialist engineering disciplines while building team capability and supporting strategic direction established by the Infrastructure Reliability Engineering Senior Manager.
- Strong understanding of Infrastructure Reliability Engineering and Site Reliability Engineering principles and practices.
- Experience defining and improving SLIs, SLOs, availability targets, and reliability metrics.
- Strong scripting/automation capability (e.g., Python, Bash, Ansible, Terraform, GitOps).
- Experience with CI/CD pipelines (GitHub Actions, Jenkins, Azure DevOps, GitLab, etc).
- Strong observability skills—including metrics, logs, distributed tracing, and tools such as Prometheus, Grafana, ELK, OpenTelemetry, Jaeger.
- Good understanding of enterprise infrastructure platforms including Linux, virtualisation, container platforms, and cloud-native technologies.
- Strong understanding of resilience engineering, disaster recovery, and service continuity principles.
- Experience conducting root cause analysis and driving long-term reliability improvements.
- Understanding of operational readiness, service transition, and supportability principles.
- Knowledge of infrastructure security, hardening, compliance controls, and operational risk management.
- Understanding of ITIL-aligned Incident, Problem, Change, and Service Management disciplines.
- Strong people leadership, coaching, and team development skills.
- Strong analytical, troubleshooting, and problem-solving capability.
- Excellent stakeholder management and relationship-building skills.
- Ability to provide technical leadership while maintaining focus on delivery, governance, and business outcomes.
- Strong planning, prioritisation, and organisational skills.
- Excellent written, verbal, and presentation skills.
- Strong focus on automation, standardisation, and continuous improvement.
- Data-driven and evidence-based approach to decision-making.
- Demonstrates ownership, accountability, and engineering excellence.
- Comfortable operating within high-pressure, business-critical, and regulated environments.
Skills
- Infrastructure Reliability Engineering
- Site Reliability Engineering
- Platform Engineering
- Infrastructure Engineering
- Observability
- Resilience Engineering
- Automation
- Service Reliability
- Operational Readiness
- SLI/SLO Definition and Improvement
- Availability Metrics
- Reliability Metrics
- Monitoring
- Telemetry
- Logging
- Tracing
- Alerting
- Service Health Monitoring
- Data-driven Analysis
- Resilience Validation
- Failover Testing
- Disaster Recovery Testing
- Recovery Assurance
- Scenario-based Resilience Assessments
- Operational Readiness Standards
- Service Continuity
- Fault Tolerance
- Infrastructure as Code
- Configuration Management
- Standardisation
- Repeatability
- Platform Reliability
- Scalability
- Recoverability
- Operational Efficiency
- Cloud-native Technologies
- Virtualisation
- Enterprise Infrastructure Platforms
- Linux
- Container Platforms
- Kubernetes
- OpenShift
- CI/CD Pipelines
- Python
- Bash
- Ansible
- Terraform
- GitOps
- Prometheus
- Grafana
- ELK
- OpenTelemetry
- Jaeger
- Root Cause Analysis
- Service Transition
- Supportability Principles
- Infrastructure Security
- Hardening
- Compliance Controls
- Operational Risk Management
- ITIL
- Incident Management
- Problem Management
- Change Management
- Service Management
- People Leadership
- Coaching
- Team Development
- Analytical Skills
- Troubleshooting
- Problem-Solving
- Stakeholder Management
- Relationship Building
- Technical Leadership
- Delivery Focus
- Governance Focus
- Business Outcomes Focus
- Planning
- Prioritisation
- Organisational Skills
- Written Communication
- Verbal Communication
- Presentation Skills
- Continuous Improvement
- Data-driven Decision-making
- Ownership
- Accountability
- Engineering Excellence
Location
- UK-London
Work Type
- Permanent
- Standard 40 Hour Week
Experience Level
- Assistant Vice President
- Significant experience in Infrastructure Reliability Engineering, Site Reliability Engineering, Platform Engineering or Infrastructure Engineering disciplines
- Demonstrable experience leading technical engineering teams within complex enterprise environments
- Significant experience operating within complex enterprise technology environments
- Strong track record of leading teams responsible for critical infrastructure services and engineering outcomes
- Extensive experience improving infrastructure reliability through engineering practices rather than operational intervention
Education Level
- Bachelor's degree in Computer Science, Engineering, Information Technology or a related technical discipline
- Professional certifications or equivalent practical experience related to Linux, cloud-native platforms, observability, automation, infrastructure engineering or reliability engineering would be advantageous
About the Company
- LME Group is the world centre for industrial metals trading and clearing.
- Most of the world’s non-ferrous metals business is conducted in the LME.
- The metals community uses the LME, a member of HKEX Group, as a venue to transfer or take on price risk, as a physical market of last resort and as the provider of transparent global reference prices.
