About the Role
TableCheck, Japan's leading restaurant reservation management platform, is seeking a Site Reliability Engineer with machine learning expertise to own the technology stack and support business and developer needs. This role focuses on maintaining and evolving production infrastructure on AWS and Kubernetes, while also contributing to machine learning initiatives by bringing reliability engineering discipline to ML pipelines, model deployment, and supporting infrastructure.
Responsibilities
- Maintain a 24/7 production environment running on Kubernetes following SRE principles.
- Implement DevOps methodologies to improve IT team quality of life.
- Perform proactive system monitoring and configuration.
- Engage in incident response and postmortem processes.
- Manage and evolve AWS infrastructure including EKS, EC2, RDS, Fargate, CloudFront, Lambda, and S3.
- Build and maintain CI/CD pipelines and infrastructure as code using Terraform, Helm, and ArgoCD.
- Ensure system reliability, performance, and scalability across the production stack.
- Apply SRE discipline to ML infrastructure, ensuring model serving, training pipelines, and data systems are reliable, observable, and well-operated.
- Support and improve ML model deployment pipelines and MLOps practices.
- Monitor ML model performance in production and build alerting and observability for ML systems.
- Collaborate with data scientists and product teams to operationalize ML models at scale.
- Contribute to infrastructure for ML workloads on Kubernetes and AWS.
Requirements
- At least 2 years of experience with Amazon Web Services (AWS), with a focus on EKS, EC2, RDS, Fargate, CloudFront, Lambda, and S3.
- Extensive hands-on experience using AWS EKS.
- At least 1 year of experience in direct software engineering following DevOps / SRE practices as a technical lead.
- Current ability in at least one of the following languages: Python, Ruby, Elixir, Go, Javascript, Rust.
- Understanding of container and hypervisor fundamentals.
- Experience with configuration management (YAML / Bash).
- Experience running production systems at large scale and understanding potential problems and solutions.
- Familiarity with machine learning workflows and MLOps practices.
- Python experience with ML-adjacent tooling (model deployment, inference serving, or ML pipeline tooling).
- Previous startup experience is highly desired.
- Experience with Terraform or Pulumi.
- Experience with ArgoCD.
- Experience with Prometheus and Grafana for monitoring and alerting.
- Experience with PostgreSQL and MongoDB.
- Experience with Kafka for event-driven architecture.
- Knowledge of Security, PCI-DSS, GDPR, and forensics.
- Experience with ML model serving frameworks (e.g., TensorFlow Serving, TorchServe, Triton).
- Familiarity with feature stores, experiment tracking, or model registry tools.
- Experience deploying and managing ML workloads on Kubernetes.
Skills
- AWS EKS
- Python
- Ruby
- Elixir
- Go
- Javascript
- Rust
- Helm
- Terraform
- ArgoCD
- Prometheus
- Grafana
- PostgreSQL
- MongoDB
- Kafka
- MLOps
- ML model serving frameworks
- Feature stores
- Experiment tracking
- Model registry tools
Location
- Remote
Work Type
- Remote
- Full-time
Experience Level
- Technical Lead
- 2+ years of AWS experience
- 1+ year as a technical lead
About the Company
- TableCheck is Japan's leading restaurant reservation management platform.
- We run a robust and fault-tolerant infrastructure built on Amazon Web Services (AWS) with Terraform, Kubernetes, Helm, and an array of tools for CI/CD, logging, monitoring, and more.
- We emphasize DevOps best practices such as agile, scrum, automation, and customer-centric improvements.
- TableCheck has embraced remote work, prioritizing communication and documentation.
- We constantly learn from mistakes and adapt, expecting team members to follow up with questions and updates.
