Join a global technology leader in Tokyo as a Senior Site Reliability Engineer (SRE / DevOps), taking ownership of the reliability, scalability, automation, and performance of high-volume transactional platforms used by millions of customers.
You will work on mission-critical rewards, points, and coupon platforms processing large volumes of transactions, partnering closely with product and development teams to define SRE best practices and improve service reliability.
This role combines Kubernetes, cloud infrastructure, Infrastructure as Code (IaC), CI/CD, observability, FinOps, and AI-driven automation. You will design highly available infrastructure, improve production operations, reduce engineering toil, and build self-service capabilities that enable development teams to deliver reliable software at scale.
Key Responsibilities
- Define, monitor, and continuously improve SLIs and SLOs in collaboration with product and software development teams.
- Design highly available and scalable infrastructure covering redundancy, disaster recovery, capacity planning, and performance optimisation.
- Build, operate, and scale production infrastructure using Linux, Kubernetes, public cloud, and Infrastructure as Code (IaC).
- Develop and maintain automated CI/CD pipelines using tools such as Jenkins and GitHub Actions.
- Lead production operations, release management, on-call processes, incident response, and blameless postmortems.
- Design and optimise core networking components, including DNS, CDN, load balancing, and TLS, for large-scale services.
- Implement observability and monitoring using technologies such as Prometheus, Grafana, ELK, and OpenTelemetry.
- Strengthen operational security through vulnerability patching, secrets management, security controls, and compliance auditing for transactional systems.
- Drive FinOps and cloud cost optimisation while balancing reliability, scalability, and system performance.
- Reduce operational toil by developing self-service tooling, golden paths, and platform engineering capabilities for application teams.
- Leverage AI coding and operations tools such as GitHub Copilot and Claude Code to automate engineering workflows and improve operational efficiency.
Required Skills and Qualifications
Experience:
- 5+ years of hands-on experience in Site Reliability Engineering (SRE), DevOps, platform engineering, or infrastructure engineering.
- 5+ years of production experience working with Linux and Kubernetes, including architecture, deployment, scaling, troubleshooting, and maintenance.
- Proven experience designing and building large-scale production infrastructure from the ground up.
- Strong knowledge of networking fundamentals, including DNS, CDN, load balancing, and TLS.
- Hands-on experience with Infrastructure as Code technologies such as Terraform or Pulumi.
- Experience building and operating CI/CD pipelines using Jenkins, GitHub Actions, or comparable technologies.
- Strong experience with observability and monitoring technologies such as Prometheus, Grafana, ELK, and OpenTelemetry.
- Production experience with at least one major public cloud platform: GCP, AWS, or Azure.
- Scripting and automation experience using Python and/or Shell.
- Practical experience using AI-assisted coding or operations tools, such as GitHub Copilot, Claude Code, or similar technologies.
- Bachelor's degree or equivalent qualification.
Soft Skills:
- Strong communication, technical writing, and stakeholder management skills, with the ability to lead complex infrastructure and reliability initiatives.
- Proactive approach to continuous improvement, automation, and challenging existing operational practices.
- Strong problem-solving skills and the ability to make effective decisions during production incidents.
- Collaborative mindset with the ability to guide, support, and empower application development teams.
- Strong ownership mentality with a focus on reliability, scalability, security, and developer productivity.
Language Requirements:
- Japanese: Basic-level proficiency. Japanese language capability is advantageous for collaboration with local stakeholders.
- English: Intermediate to business-level proficiency required for technical communication and collaboration within an international engineering environment.
Preferred Skills & Qualifications
- Production experience operating stateful middleware and databases, including RDBMS, NoSQL databases, Redis, and message queues.
- Familiarity with modern Platform Engineering concepts, including Team Topologies, DORA metrics, Internal Developer Platforms (IDPs), and golden paths.
- Experience designing or operating large-scale, high-concurrency platforms supporting 100,000+ users.
- Previous experience as a software or application engineer.
- Experience designing self-service infrastructure and developer platforms that reduce cognitive load for engineering teams.
- Knowledge of FinOps, cloud cost management, disaster recovery, and capacity planning.
- Experience supporting high-volume transactional, payment, rewards, e-commerce, or other mission-critical digital platforms.
- Japanese language proficiency beyond the basic level.
About the Company
Our client is a global technology conglomerate and digital platform leader operating an interconnected ecosystem of services used by millions of customers every day.
As the organisation continues to scale its mission-critical points, rewards, and coupon platforms, reliability and platform engineering have become increasingly important to delivering fast, resilient, and highly available digital services.
The engineering environment makes extensive use of cloud-native technologies, Kubernetes, Infrastructure as Code, observability, CI/CD, and AI-driven automation, giving SRE and DevOps engineers the opportunity to solve complex infrastructure challenges at significant scale.
Why You'll Love Working Here
- Take technical ownership of large-scale transactional platforms processing extremely high volumes of operations.
- Work extensively with Kubernetes, cloud infrastructure, Terraform, CI/CD, observability, and platform engineering.
- Gain practical exposure to AI-driven engineering and operations automation, including modern AI coding tools.
- Influence critical architectural decisions across high availability, disaster recovery, capacity planning, reliability, and FinOps.
- Join an engineering culture focused on reducing toil, developer enablement, automation, and blameless postmortems.
- Collaborate with an international engineering team in a highly technical environment.
- Build self-service platforms and golden paths that improve productivity across application development teams.
- Benefit from remote work/WFH options, flex time, minimal overtime, and casual clothing.
Don't Miss Out - Apply Now!
