About the Role
As a Site Reliability Engineer (SRE) on the Cloud Platform team, you will shape the reliability, scalability, and performance of our Cloud platform and customer-facing applications. This role is critical in maintaining the stability and efficiency of our infrastructure, enabling seamless experiences for users and developers.
Responsibilities
- Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support our Cloud platform.
- Operate systems and troubleshoot issues in production environments, including interrupts, on-call responses, and infrastructure scaling.
- Implement and improve monitoring, alerting, and incident response systems to minimize downtime and optimize performance.
- Develop and maintain workflows and tools for CI/CD, containerization, orchestration, monitoring, and logging.
- Participate in on-call rotations to respond to incidents and perform root cause analysis to prevent recurrence.
- Drive continuous improvement in infrastructure automation, deployment, and orchestration.
- Collaborate with software engineers to enable safe and reproducible model-training experiments.
- Build and enhance a cloud platform that abstracts infrastructure complexities for science and engineering teams.
- Design and develop new workflows, tooling, and automation to improve system reliability, availability, and performance.
- Ensure infrastructure adheres to security best practices and compliance requirements in collaboration with the security team.
- Document processes and procedures to ensure consistency and knowledge sharing across the team.
Requirements
- 5+ years of experience in a DevOps or SRE role, with strong expertise in bare metal infrastructure and distributed systems.
- Hands-on experience with site reliability issues, including root cause analysis, in-production troubleshooting, and on-call rotations.
- Proficiency in working with reliability KPIs, such as observability, alerting, and SLAs.
- Experience with CI/CD, containerization, and orchestration tools like Docker and Kubernetes.
- Knowledge of monitoring, logging, alerting, and observability tools such as Prometheus, Grafana, ELK Stack, or Datadog.
- Familiarity with infrastructure-as-code tools like Terraform or CloudFormation.
- Proficiency in scripting languages (Python, Go, Bash) and a strong understanding of software development best practices.
- Solid grasp of networking, security, and system administration concepts.
- Excellent problem-solving and communication skills, with the ability to work effectively in a collaborative environment.
- Experience in an AI/ML environment, high-performance computing (HPC) systems, or modern AI-oriented solutions (e.g., Fluidstack, Coreweave, Vast) is a plus.
Skills
- Docker
- Kubernetes
- Prometheus
- Grafana
- ELK Stack
- Datadog
- Terraform
- CloudFormation
- Python
- Go
- Bash
Experience Level
- 5+ years
Education Level
- Master’s degree in Computer Science, Engineering, or a related field
Benefits
- Healthcare coverage
- Parental leave
- Retirement plans
- Relocation support
- Wellness programs
- Meal and transportation allowances
About the Company
- Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute.
- We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.
- We are a dynamic, collaborative team passionate about AI and its potential to transform society.
- Our diverse workforce thrives in competitive environments and is committed to driving innovation.
- Our teams are distributed between Europe, North America, Asia and the Middle East.
- We are creative, low-ego and team-spirited.
