RemoteFull timeMid levelPosted today
Apply with JobAssistAbout the role
- Our Site Reliability Engineering (SRE) team consists of highly skilled engineers responsible for maintaining and enhancing the reliability, scalability, and performance of the ServiceNow infrastructure
- Our SRE’s are empowered to resolve technical issues across the entire technology stack, from hardware to applications
- Additionally, they work to improve the platform’s operability, aiming to reduce the number of incidents and minimize Mean Time to Recovery (MTTR)
- To achieve this, the team combines software development, networking, database, and systems engineering skills to tackle complex problems, striving to maintain our platform operating for our customers
- As a Sr Manager, Site Reliability Engineering at ServiceNow, you’ll lead a team of SRE leaders, managers, and engineers focused on ensuring the reliability and availability of critical enterprise platforms and applications, while helping drive our cloud modernization journey through automation, operational excellence, resilience, and continuous improvement
- Lead and develop a global team of SRE leaders, managers, and engineers, with accountability for talent development, performance, prioritization, succession planning, and execution
- Define and drive the SRE strategy and operating model across reliability, observability, automation, incident response, production readiness, and continuous improvement
- Own and evolve observability capabilities across metrics, logs, traces, alerting, SLI/SLOs, error budgets, and service health practices to improve detection, diagnosis, and overall production reliability
- Lead the strategy and execution for production-like staging environments, ensuring critical services and releases can be validated in environments that closely represent production
- Partner with Engineering and Release teams to strengthen release pipelines and confidence gates for Now releases, including smoke, integration, resiliency, performance, rollback, and production-readiness validation
- Drive a culture of eliminating repetitive operational work through automation, orchestration, self-healing, and AI-assisted operations, shifting reliability practices earlier in the software-development lifecycle
- Lead cloud modernization initiatives, modernizing legacy infrastructure, tooling, and operational practices toward cloud-native architectures and scalable engineering patterns
- Drive adoption and operational excellence across AWS, GCP, and Azure hyperscalers, including architecture, scaling, resiliency, availability, cost efficiency, and operational readiness
- Advance containerization and Kubernetes-based operating models, including workload resiliency, scalability, orchestration, deployment patterns, observability, and lifecycle management
- Partner with Product Engineering, Platform Engineering, Release Engineering, Security, and Infrastructure teams to improve the reliability and availability of critical enterprise platforms and services
- Establish reliability standards and measurable outcomes around SLIs/SLOs, error budgets, availability, MTTR, change failure rate, automation, and operational readiness
- Provide leadership during major incidents, while ensuring incident and problem-management learnings are translated into durable engineering improvements, automation, and preventative controls
- Establish follow-the-sun and on-call SRE practices that keep engineers close to production signals and use operational learnings to drive shift-left improvements
- Evaluate existing platforms, processes, and technologies and drive simplification, modernization, standardization, and operational efficiency
- Identify and adopt emerging technologies that support the organization’s reliability and cloud-modernization goals, while building the skills and capabilities required to operate them at scale
- Build a global engineering culture that values technical excellence, accountability, collaboration, continuous learning, and diverse perspectives
Benefits
- Generous family leave
- Matched donations
- Annual learning stipends
- Flexible PTO
- Competitive retirement plan
- Paid volunteer time- Know operating systems in various levels of troubleshooting and diagnostics?
- Have experience in leading a team of engineers and exposure to people management?
- Have experience with cloud technologies and hyperscalers such as AWS, GCP, or Azure, and a passion for driving cloud modernization and cloud-native transformation?
- Have a technical background in roles like systems engineering or devops or site reliability engineering?
- Have low tolerance to repetitive tasks and automate your way through work?
- If you Answered ‘yes’ to these questions, we want to hear from you. Hit the Apply button and let’s have a chat about the role and your skills and experiences
- Excellent written and verbal communication skills, with the ability to translate complex technical topics into clear business outcomes and leadership decisions
- Experience applying automation, orchestration, and infrastructure-as-code to reduce operational toil and improve repeatability and reliability
- Demonstrated ability to influence organizational boundaries and drive alignment among Engineering, Product, Architecture, Security, Infrastructure, and Operations teams
- Strong technical foundation across Linux, distributed systems, databases, networking, systems troubleshooting, scripting, and software engineering fundamentals
- Experience operating high-scale software, platform, and infrastructure-as-a-service environments with demanding availability and reliability requirements
- 5+ years of people-management experience, including experience leading managers, senior technical leaders, and geographically distributed engineering teams
- Experience building or operating production-like staging/test environments and establishing validation strategies for reliability, resiliency, performance, integration, and release readiness
- Experience designing and operating observability platforms at scale, including metrics, logging, tracing, alerting, dashboards, golden signals, SLI/SLOs, and error budgets
- Strong understanding of incident management, problem management, operational readiness, and continuous improvement practices
- Deep working knowledge of one or more major hyperscalers - AWS, GCP, or Azure with experience designing and operating resilient, scalable, highly available production systems
- Experience leveraging or critically evaluating AI and AI-assisted operations to automate workflows, accelerate diagnosis and remediation, and improve engineering productivity
- Experience with CI/CD and release pipelines, including automated confidence gates, progressive/phased deployments, zero-downtime deployment strategies, rollback mechanisms, and production-readiness controls
- Strong experience with cloud modernization and migration, including modernizing legacy platforms and tooling into cloud-native architectures
- Significant experience leading Site Reliability Engineering, Production Engineering, DevOps, Platform Engineering, or Cloud Infrastructure organizations in large-scale production environments
- Strong understanding of Kubernetes, containers, orchestration, service networking, autoscaling, and modern cloud-native architecture patterns
- Ability to lead effectively through ambiguity and change while maintaining a strong focus on execution, customer impact, and engineering excellence
- RHCE, CCNA, ITIL or other industry certifications
Millions of jobs, with real people getting hired every day
20,000+
New jobs added daily7,000,000+
Verified job listings500,000+
Tailored applications submittedFAQ
Questions, answered
Click "Apply with JobAssist" – we tailor your resume and application to this role and submit it for your approval.
Yes. This role at ServiceNow was screened before publishing – we confirmed the employer before listing it.
The employer didn't disclose a salary range for this listing. JobAssist shows pay whenever it's available.
This position can be done from anywhere, with no in-office requirement.
Yes – every application is tailored from your profile and this job's requirements, and you can review and edit before it's sent.
