IonQ logo

Staff Site Reliability Engineer

IonQ

RemoteFull timeMid level$163k – $214kPosted today
Apply with JobAssist

About the role

Who you are

  • 7+ years of production engineering experience with recent hands-on reliability work
  • Hands-on, recent experience operating large-scale, fault-tolerant production systems on AWS or GCP
  • Observability ownership: have instrumented production systems and governed service-level objectives and error budgets, not only installed dashboards
  • Resilience practice: have designed and executed failure experiments or disaster-recovery exercises with real failover validation
  • Incident command: have personally commanded serious SEV1/SEV2 incidents and driven root cause through to a systemic fix
  • Demonstrated ownership of reliability outcomes with measurable results, such as availability, mean time to recovery, and error-budget adherence
  • Evidence of multi-team technical leadership through standards, review, coaching, and mechanisms adopted beyond one service or team
  • Proven production experience with cloud security posture management, runtime vulnerability detection, and workload protection across cloud and distributed environments
  • Strong experience prioritizing risk using identity, workload, and exposure-path context to focus remediation on issues that materially increase attack likelihood and operational impact
  • Experience with autonomous remediation and self-healing workflows powered by AIOps, including Amazon Bedrock Agent Core or equivalent agentic automation frameworks
  • Hands-on experience in capacity management, resource rightsizing, efficiency engineering, and practical cost optimization based on FinOps principles
  • Experience with load-balancing design and operations, including health-based failover, global traffic management, and performance optimization for highly available services
  • Experience with AI traffic management via an LLM gateway, including request routing, policy enforcement, rate limiting, model fallback, latency optimization, cost controls, and observability for multi-model or multi-provider environments
  • Ability to connect networking, security, and reliability considerations into cohesive platform design decisions that improve resilience, performance, and operability

What the job involves

  • We are seeking a Staff Site Reliability Engineer. As Staff SRE Engineer, you set the technical direction for reliability across regions and services
  • You own the reliability strategy, define the standards and mechanisms that guide production operations, and raise the bar through design leadership, operational discipline, and mentorship
  • You remain deeply hands-on by designing and operating observability platforms, defining and governing SLO programs, leading high-severity incident response, building resilience and disaster-recovery automation, improving reliability of stateful and streaming platforms, and creating AI Ops workflows for triage, remediation, and self-healing
  • Production reliability: own service-level objectives, error budgets, and production reliability outcomes end to end, and represent reliability in architecture and scaling decisions
  • Engineer observability: design and operate the observability stack so production services are fully instrumented and define the standards platform and application teams follow
  • Govern SLOs and error budgets: define and manage service-level objectives, run regular reviews with service owners, and drive corrective action when services consume error budgets unsafely
  • Drive resilience: design and execute chaos experiments and validate that failure modes are covered by tested safeguards
  • Lead incident response: define the incident process and serve as incident commander for the highest-severity incidents, including security incidents within the coverage window
  • Run on-call and escalation: establish and manage rotations and escalation paths that provide continuous coverage with clean follow-the-sun handoffs
  • Disaster recovery: own disaster-recovery testing and failover validation against defined recovery objectives and turn exercise findings into architectural and operational improvements
  • Cloud security posture: co-own cloud security posture management, runtime vulnerability detection, and configuration-compliance monitoring with DevSecOps
  • Data, streaming, and AI Ops: own reliability of stateful and streaming services, capacity planning and rightsizing, and autonomous agents for triage, predictive alerting, remediation, and self-healing
  • Scale the team and broaden impact: mentor engineers at different seniority levels, set standards adopted across teams, and align Architecture, DevSecOps, Cloud Operations, and Product Development behind a shared reliability roadmap

Millions of jobs, with real people getting hired every day

20,000+
New jobs added daily
7,000,000+
Verified job listings
500,000+
Tailored applications submitted
FAQ

Questions, answered

Click "Apply with JobAssist" – we tailor your resume and application to this role and submit it for your approval.

Yes. This role at IonQ was screened before publishing – we confirmed the employer before listing it.

The employer didn't disclose a salary range for this listing. JobAssist shows pay whenever it's available.

This position can be done from anywhere, with no in-office requirement.

Yes – every application is tailored from your profile and this job's requirements, and you can review and edit before it's sent.