ML Operations Engineer (AI/LLM) - Mercari at Mercari, inc. | JP | Rezi

ML Operations Engineer (AI/LLM) - Mercari at Mercari, inc.

ML Operations Engineer (AI/LLM) - Mercari

Mercari, inc. · JP

Yesterday

ML Operations Engineer (AI/LLM) - Mercari

Mercari, inc. · JP

a day ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

As an MLOps engineer on the AI/LLM team, you will own how our machine learning and LLM models reach production and stay healthy in our cloud-native environment. Your focus is the production serving, deployment, and operations that turn models into reliable, cost-efficient services, seamlessly integrating with our machine learning operations to serve tens of millions of users.

Responsibilities

  • Own the end-to-end orchestration of model inference, including integrating with DataServ for retrieval (BigQuery, BigTable, Valkey) and managing the Model Inference Gateway and Console.
  • Own production model serving on the cloud-native NVIDIA and TPU stacks (Triton Inference Server, TensorRT-LLM, JAX/TPU Gateways).
  • Manage model repositories, dynamic batching, and concurrent model execution.
  • Build CI/CD, rollout, and rollback paths for safe, scalable model deployment.
  • Automate provisioning and lifecycle management (Terraform, Kubernetes) so the platform scales seamlessly across teams.
  • Profile and optimize deployments across LLM and non-LLM workloads, including model compilation, quantization, and batching strategies to hit latency and throughput targets while managing cost.
  • Maintain performance baselines and regression detection.
  • Build robust monitoring and alerting for model health and latency, including service-level metrics for the data retrieval and inference gateway layers.
  • Define SLOs and own on-call and incident response for the serving layer.
  • Build automated evaluation and quality monitoring into the deployment path, regression and drift detection, offline/online evaluation, and LLM output quality checks, so models stay healthy long after launch.
  • Partner with ML engineers and researchers to turn experimental models into production-ready services, providing self-service workflows and abstractions that let teams deploy safely and quickly without deep infrastructure expertise.

Requirements

  • Shared belief in the mission and values of Mercari Group.
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 5+ years of software engineering experience, including proven experience in production MLOps: end-to-end model deployment, serving, and CI/CD in cloud environments.
  • Experience designing and operating large-scale, high-availability distributed systems, including observability, SLO definition, and incident response.
  • Strong experience in cloud-native infrastructure (Kubernetes, Docker).
  • Proficiency in Python and infrastructure-as-code (Terraform).
  • Excellent written and verbal communication.
  • Experience integrating ML serving with large-scale distributed data layers (e.g., data warehouses, wide-column stores, in-memory caches).
  • Expertise in model inference optimization (TensorRT-LLM, quantization, JAX).
  • Experience operating large-scale model inference gateways and orchestrators.
  • 2+ years of hands-on experience operating GenAI/LLM workloads in production (e.g., LLM serving frameworks, token throughput and cost optimization).
  • Experience building LLM evaluation, guardrail, or quality-monitoring pipelines (e.g., LLM-as-judge, golden datasets, drift detection).
  • Experience with serving infrastructure for RAG or agentic AI workloads (vector search, tool-calling execution environments).
  • Experience partnering closely with research or data science teams to bring research innovations into production.
  • Master's or Ph.D. in a related technical field.
  • English: Proficient (CEFR - B2)
  • Japanese: Independent (CEFR - B2) optional

Skills

  • MLOps
  • Model Deployment
  • Model Serving
  • CI/CD
  • Cloud Environments
  • Kubernetes
  • Docker
  • Python
  • Terraform
  • NVIDIA
  • TPU
  • Triton Inference Server
  • TensorRT-LLM
  • JAX
  • Data Orchestration
  • Model Inference Gateway
  • Model Repository Management
  • Dynamic Batching
  • Concurrent Model Execution
  • Rollout
  • Rollback
  • Provisioning
  • Lifecycle Management
  • Performance Profiling
  • Model Compilation
  • Quantization
  • Latency Optimization
  • Throughput Optimization
  • Cost Management
  • Performance Baselines
  • Regression Detection
  • Monitoring
  • Alerting
  • Service-Level Metrics
  • SLOs
  • On-call
  • Incident Response
  • Model Quality
  • Automated Evaluation
  • Regression Detection
  • Drift Detection
  • Offline Evaluation
  • Online Evaluation
  • LLM Output Quality Checks
  • Research-to-Production Enablement
  • Self-service Workflows
  • Distributed Systems
  • Observability
  • RAG
  • Agentic AI Workloads
  • Vector Search
  • Tool-calling Execution Environments

Location

  • Roppongi

Work Type

  • Full-time
  • Full Flextime (no core time)

Experience Level

  • 5+ years of software engineering experience
  • 2+ years of hands-on experience operating GenAI/LLM workloads in production

Education Level

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • Master's or Ph.D. in a related technical field.

About the Company

  • Circulate all forms of value to unleash the potential in all people.
  • Mercari marketplace app was born in 2013 out of this thought by our founder Shintaro Yamada as he traveled the world.
  • We believe that by circulating all forms of value, not just physical things and money, we can create opportunities for anyone to realize their dreams and contribute to society and the people around them.
  • Mercari aims to use technology to connect people all over the world and create a world where anyone can unleash their potential.
  • Mercari Engineering Principles are a shared understanding that serves as the foundation of engineering beliefs and behavior at Mercari.
  • The Engineering Principles are designed to complement the organizational identity (Mercari’s mission, values, and culture) from an engineering viewpoint.
  • These principles ultimately help us achieve Mercari’s mission by defining the ideal state we seek to realize in the long term.
  • Passion For The Product
  • Grow Together
  • Solve Through Mechanisms
  • Collaborate Openly
  • The AI / LLM Team’s mission is focused on three core pillars, “product”, "enablement" and "research", delivering new AI-driven features and user experiences to maximize product-facing impact for Mercari's business.
  • We do this both through independent initiatives owned by our team, as well as by horizontally collaborating with product, engineering, and research teams across the entire organization.

Equal Opportunity

  • Here at Mercari, we work to realize a world in which no one’s potential is limited by their background and everyone has the opportunity to freely create value.
  • We also firmly believe that a mindset of Inclusion & Diversity is essential for us to achieve our mission.
  • Mercari is committed to eliminating discrimination based on age, gender, sexual orientation, race, religion, physical disability, and other such factors so that anyone who shares our mission and values can join us, regardless of their background.