Site Reliability Engineer
Standard baseline compensation
Verified employer route
Paris
100% direct official career link
“Mistral.ai Verified Licensed Visa Sponsor (GB)”
Find More Opportunities With AI Smart Match
Match your profile against 12,000+ verified visa sponsorship roles across UK, USA, Australia, and Canada.
Job Description & Specifications
Official role breakdown, eligibility standards, and core responsibilities
Site Reliability Engineer at Mistral.ai
📍 Paris
What You Will Do
- Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support web services and ML workloads.
- Ensure our platform, inference, and model training environments are always highly available and enable seamless replication across HPC clusters.
- Operate systems and troubleshoot issues in production, including interrupts, on-call responses, and infrastructure scaling.
- Implement and improve monitoring, alerting, and incident response systems to minimize downtime and optimize performance.
- Develop and maintain workflows and tools for CI/CD, containerization, orchestration, monitoring, and logging.
- Participate in on-call rotations to respond to incidents and perform root cause analysis.
- Drive continuous improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, and Terraform.
- Collaborate with AI/ML researchers to enable safe and reproducible model-training experiments.
- Build a cloud-agnostic platform that abstracts infrastructure complexities for science and engineering teams.
- Design and develop new workflows, tooling, and automation to improve system reliability, availability, and performance.
- Work with the security team to ensure infrastructure adheres to best practices and compliance requirements.
- Document processes and procedures to ensure consistency and knowledge sharing across the team.
What We're Looking For
- A Master’s degree in Computer Science, Engineering, or a related field.
- 7+ years of experience in a DevOps or SRE role, with strong expertise in cloud computing and distributed systems.
- Hands-on experience with site reliability issues, including root cause analysis, in-production troubleshooting, and on-call rotations.
- Proficiency in working with reliability KPIs, such as observability, alerting, and SLAs.
- Experience with CI/CD, containerization, and orchestration tools like Docker and Kubernetes.
- Knowledge of monitoring, logging, alerting, and observability tools such as Prometheus, Grafana, ELK Stack, or Datadog.
- Familiarity with infrastructure-as-code tools like Terraform or CloudFormation.
- Proficiency in scripting languages (Python, Go, Bash) and a strong understanding of software development best practices.
- Solid grasp of networking, security, and system administration concepts.
- Excellent problem-solving and communication skills, with the ability to work effectively in a collaborative environment.
- Experience in an AI/ML environment, high-performance computing (HPC) systems, or modern AI-oriented solutions (e.g., Fluidstack, Coreweave, Vast) is a plus.
Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.Find Jobs in France on Arbeitnow
Required Skills / Keywords: Engineering & Infra
Application Process: Click "Apply for this Job" to be directed to the employer's official application page.
Verified Sponsor Job — VIP Only
This role is on the verified Official Register of Licensed Sponsors. Unlock the full description, salary package, and direct application link.
🔒 One-time payment via Razorpay. No auto-renewals. Instant unlock.
Start Your Application
Job Overview
AI Career Intelligence Match
Discover adjacent roles and verified visa-sponsored positions matching your background.
Similar Verified Sponsor Vacancies
Explore other roles offering visa sponsorship in United Kingdom