About the Role
As the engineering organization transitions from a monolithic architecture to a service-oriented architecture, this role centers on transforming core environments into an advanced, self-serve developer platform. You will build high-level, composable abstractions that automate pipelines, infrastructure, logging, and metrics, allowing product engineers to deploy production services rapidly while safely hiding underlying complexities.
Responsibilities
- Collaborate with the infrastructure group to build a self-serve platform that enables engineering teams to scale services from inception to production deployment via composable, high-level abstractions.
- Treat internal software engineers as your primary customers. Gather direct feedback, isolate operational friction points, and design foundational architectures that teams naturally adopt rather than bypass.
- Actively evaluate, benchmark, and determine the components of the core platform stack, including container orchestration models, end-to-end observability tools, and defining clear boundaries for buy-versus-build technical decisions.
- Establish strict institutional benchmarks for infrastructure reliability, systemic visibility, and propagate an engineering-led "you build it, you run it" operational culture.
- Partner directly with cross-functional development groups to facilitate smooth application migrations out of legacy systems into distributed microservices.
- Raise technical execution standards by mentoring engineers, performing comprehensive design and code reviews, and driving technical decisions that safeguard multi-year scalability.
Requirements
- Proven experience designing and implementing internal developer platforms or specialized internal tooling that driving widespread engineering adoption.
- Deep production-level containerization experience with Kubernetes systems, specifically handling enterprise-grade networking, RBAC policies, auto-scaling mechanisms, and custom controllers.
- Strong mastery of Terraform, including writing reusable and modular code configurations.
- Proficient in architecting scalable pipelines using modern continuous delivery systems (e.g., DroneCI, ArgoCD, GitHub Actions).
- Direct, hands-on architectural experience managing core cloud ecosystems, including advanced identity access management (IAM), virtual private networks (VPC), and managed database endpoints (GCP experience is a welcome addition).
- Proven depth optimizing comprehensive telemetry infrastructures across enterprise logging, metric aggregation, and distributed tracing via standard tooling (e.g., Grafana, Prometheus, OpenTelemetry).
- Advanced knowledge of platform-level networking and security models, including secrets management implementations, secure service-to-service authentication protocols, and software supply chain protection.
- Comfortable maintaining and provisioning stateful relational and non-relational database storage environments in production setups (e.g., RDS, Aurora, MySQL, MongoDB, Redis, Elasticsearch).
- Practical, real-world experience managing structured incident response procedures, participating in production on-call rotations, running blameless post-mortems, and embedding structural cost awareness directly into architecture designs.
- Exceptional written and verbal English communication skills.
- Superior aptitude for decomposing complex, highly ambiguous, and ill-defined technical problems into execution phases without requiring constant structural direction.
- Strong capacity to mentor peers, define explicit project task frameworks, deliver actionable and constructive feedback, and constructively influence complex architectural strategy across multi-disciplinary engineering teams.
- A proactive, entrepreneurial approach to system reliability combined with an acute sense of engineering ownership, urgency, and an understanding of how technical directions interface with broader corporate goals.
Skills
- Kubernetes
- Terraform
- DroneCI
- ArgoCD
- GitHub Actions
- GCP
- Grafana
- Prometheus
- OpenTelemetry
- RDS
- Aurora
- MySQL
- MongoDB
- Redis
- Elasticsearch
Location
- Tokyo
Work Type
- Hybrid
Benefits
- Complete flextime system with no core hours enforced.
- Option to work from home, with a requirement to commute to the office twice a week.
- Complete weekends (Saturdays and Sundays) and national holidays off.
- 10 days of annual paid vacation provided immediately at the initial time of employment.
- 6 days of designated Year-end and New Year vacations (December 29 to January 3).
- Full commuter allowance provided.
- Comprehensive social insurance enrollment, comprising complete employment insurance, workers' compensation insurance, health insurance, and welfare pension plans.
