About the Role
Mistral is seeking highly experienced Site Reliability Engineers (SRE) to shape the reliability, scalability, and performance of our platform and customer-facing applications. You will work closely with software engineers and research teams to ensure our systems meet and exceed internal and external customer expectations.
Responsibilities
- Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support web services and ML workloads.
- Ensure platform, inference, and model training environments are highly available and enable seamless replication of work environments across HPC clusters.
- Operate systems and troubleshoot issues in production environments, including interrupts, on-call responses, user administration, data extraction, and infrastructure scaling.
- Implement and improve monitoring, alerting, and incident response systems to ensure optimal system performance and minimize downtime.
- Implement and maintain workflows and tools (CI/CD, containerization, orchestration, monitoring, logging, and alerting systems) for client-facing APIs and large training runs.
- Participate in on-call rotations to respond to incidents and perform root cause analysis.
- Drive continuous improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, and Terraform.
- Collaborate with AI/ML researchers to develop and implement solutions for safe and reproducible model-training experiments.
- Build a cloud-agnostic platform offering an abstraction layer between science and infrastructure.
- Design and develop new workflows and tooling to improve system reliability, availability, and performance.
- Collaborate with the security team to ensure infrastructure adheres to best security practices and compliance requirements.
- Document processes and procedures for consistency and knowledge sharing.
- Contribute to open-source projects, research publications, blog articles, and conferences.
Requirements
- Master’s degree in Computer Science, Engineering, or a related field.
- 7+ years of experience in a DevOps/SRE role.
- Strong experience with cloud computing and highly available distributed systems.
- Exposure to site reliability issues in critical environments (issue root cause analysis, in-production troubleshooting, on-call rotations).
- Experience working against reliability KPIs (observability, alerting, SLAs).
- Hands-on experience with CI/CD, containerization, and orchestration tools (Docker, Kubernetes).
- Knowledge of monitoring, logging, alerting, and observability tools (Prometheus, Grafana, ELK Stack, Datadog).
- Familiarity with infrastructure-as-code tools like Terraform or CloudFormation.
- Proficiency in scripting languages (Python, Go, Bash) and knowledge of software development best practices.
- Strong understanding of networking, security, and system administration concepts.
- Excellent problem-solving and communication skills.
- Self-motivated and able to work well in a fast-paced startup environment.
- Experience in an AI/ML environment.
- Experience with high-performance computing (HPC) systems and workload managers (Slurm).
- Experience with modern AI-oriented solutions (Fluidstack, Coreweave, Vast).
Skills
- Kubernetes
- Flux
- Terraform
- Docker
- Prometheus
- Grafana
- ELK Stack
- Datadog
- Python
- Go
- Bash
- Cloud computing
- Distributed systems
- CI/CD
- Containerization
- Orchestration
- Monitoring
- Logging
- Alerting
- Observability
- Infrastructure-as-code
- Networking
- Security
- System administration
- AI/ML
- HPC
- Slurm
Experience Level
- 7+ years of experience
Education Level
- Master’s degree in Computer Science, Engineering or a related field
Benefits
- Comprehensive benefits package designed to support well-being, growth, and work-life balance.
- Healthcare coverage
- Parental leave
- Retirement plans
- Relocation support
- Wellness programs
- Meal and transportation allowances
- Other location-specific perks
About the Company
- Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute.
- Partners with enterprises tackling the hardest problems across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector.
- Co-creates customized AI systems that enterprises can run on their terms.
- Dynamic, collaborative team passionate about AI and its potential to transform society.
- Diverse workforce thrives in competitive environments and is committed to driving innovation.
- Teams are distributed between Europe, North America, Asia, and the Middle East.
- Creative, low-ego, and team-spirited culture.
