We are looking for a Senior Site Reliability Engineer who can help us solve problems and enhance our capabilities by supporting applications, services, and platforms. This position is to support the AI Data and User Experience program which will establish a Unified AI, Data & UX Organization to enable faster delivery, differentiated AI Products, and scalable innovation.
Tech Skills:
Unix, Shell Scripting, SQL, Python, Apache Nifi, Splunk, Dynatrace, Jenkins, GIT, XLR, AI and Agentic workflows etc.
This role provides the opportunity to influence and operate the foundational AI platforms that enable innovation. Rather than focusing solely on individual AI models or applications, you will help build and scale the enterprise platforms, tooling, operational practices, and cloud infrastructure that support the next generation of AI capabilities across the organization.
Role:
Business Operations is leading the DevOps transformation through our tooling and by being an advocate for change & standards throughout the development, quality, release, and product organizations. We need team members with an appetite for change and pushing the boundaries of what can be done with automation. Experience in working across development, operations, and product teams to prioritize needs and to build relationships is a must.
Engage in and improve the whole lifecycle of services from inception and design, through deployment, operation and refinement.
Analyse ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns
Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.
Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
Scale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve reliability and velocity.
Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead in DevOps automation and best practices.
Practice sustainable incident response and blameless post-mortems.
Take a holistic approach to problem solving, by connecting the dots during a production event thru the various technology stack that makes up the platform, to optimize mean time to recover
Work with a global team spread across tech hubs in multiple geographies and time zones
Share knowledge and mentor junior resources
Qualifications:
BS degree in Computer Science or related technical field involving coding (e.g., physics or mathematics), or equivalent practical experience.
Experience with algorithms, data structures, scripting, pipeline management, and software design.
Systematic problem-solving approach, coupled with strong communication skills and a sense of ownership and drive.
Ability to help debug and optimize code and automate routine tasks.
We support many different stakeholders. Experience in dealing with difficult situations and making decisions with a sense of urgency is needed.
Experience in one or more of the following is preferred: C, C++, Java, Python, Go, Perl or Ruby.
Experience supporting production AI, machine learning, data platforms, or large-scale distributed systems.
We need team members with an appetite for change and pushing the boundaries of what can be done with automation. Experience in working across development, operations, and product teams to prioritize needs and to build relationships is a must.
Experience in industry standard CI/CD tools like Git/BitBucket, Jenkins, Maven, Artifactory, and Chef. Experience designing and implementing an effective and efficient CI/CD flow that gets code from dev to prod with high quality and minimal manual effort is desired.
