About the Role
We are seeking a Principal Platform Engineer to serve as a senior technical anchor within our Platform/Operations function. This role is 70% hands-on engineering, focusing on designing, building, and operating the infrastructure for a global trading platform, while also providing technical mentorship and design authority to the team. A key aspect involves leading the adoption of AI in operations, utilizing LLM-based tooling and automation to reduce toil, enhance incident response, and increase operational efficiency.
Responsibilities
- Design, build, and operate infrastructure across AWS (VPC, networking, EKS) and Azure (AKS, AKV, networking), with an understanding of on-prem environments.
- Own core networking design and implementation within cloud environments, including connectivity, routing, DNS, load balancing, firewalls, and private links.
- Engineer for high availability, focusing on multi-region and multi-AZ architectures, failover design, capacity planning, and disaster recovery.
- Build and maintain infrastructure-as-code (Terraform or similar), CI/CD pipelines, and Kubernetes platforms.
- Define and drive SLOs, error budgets, and observability standards across the platform.
- Participate in post-incident reviews and drive resulting reliability work to completion.
- Continuously reduce toil through automation.
- Identify, prototype, and productionize AI-assisted workflows for incident triage, runbook automation, log/alert analysis, change-risk assessment, and internal knowledge tooling.
- Utilize AI-assisted engineering tools (e.g., Copilot) and coach the team on their safe and effective use.
- Establish guardrails for AI use in a regulated, availability-critical environment.
- Act as a mentor to engineers across Platform and Operations teams, improving design, code, and operational practices.
- Serve as a design authority on cross-team projects, reviewing architectures and ensuring new services are operable, observable, and resilient.
- Contribute to the technical roadmap for the platform function, balancing reliability investment with delivery.
Requirements
- Background in Operations or SRE running highly available, redundant production platforms.
- Deep hands-on experience with AWS (VPC design, networking, EKS) and Azure (AKS, Key Vault, networking).
- Strong Networking fundamentals: TCP/IP, routing concepts, firewalls, load balancing, hybrid connectivity (Direct Connect / ExpressRoute, VPNs).
- Production Kubernetes experience at scale, including day-2 operations.
- Proficiency in Infrastructure-as-code and automation (Terraform, Ansible, or similar; strong scripting in Python, Go, or Bash).
- Demonstrable, practical use of AI tooling to improve engineering or operational workflows.
- Credibility and communication skills to mentor senior engineers and influence design decisions.
- Understanding of different database technologies.
Skills
- AWS
- Azure
- Networking
- Kubernetes
- Terraform
- Ansible
- Python
- Go
- Bash
- AI tooling
- LLM-based tooling
- Agentic workflows
- Automation
- Infrastructure-as-code
- CI/CD
- Observability
- SLOs
- Error budgets
- Database technologies
Location
- Global
Work Type
- Hands-on engineering
- Full-time
Experience Level
- Principal
- Senior technical
About the Company
- Trayport is committed to creating and sustaining a collegial work environment in which all individuals are treated with dignity and respect and one which reflects the diversity of the community in which we operate.
- We provide accommodations for applicants and employees who require it.
