Senior Staff Service Reliability and Operational Intelligence Engineer
at IonQ
- Seniority
- Staff Principal
- Location
- Santa Clara, California, United States
- Posted
- 5d ago
at IonQ
<div class="content-intro"><p><strong>About IonQ: </strong><strong><br></strong></p> <p><u><a href="https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fionq.com%2F&esheet=54451094&newsitemid=20260316477473&lan=en-US&anchor=IonQ%2C+Inc&index=7&md5=0818ffbedc6dbad8c4d1611ac94c2a0f" target="_blank">IonQ, Inc</a></u>. [NYSE: IONQ] is the world’s leading quantum platform and merchant supplier - delivering integrated quantum solutions across computing, networking, sensing, and security. IonQ’s newest generation of quantum computers, the IonQ Tempo, is the latest in a line of cutting-edge systems that have been helping customers and partners including <u><a href="https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fionq.com%2Fnews%2Fionq-speeds-quantum-accelerated-drug-development-application-with&esheet=54451094&newsitemid=20260316477473&lan=en-US&anchor=Amazon+Web+Services%2C&index=8&md5=a7f94ee9b71a392dca3cd6c413722e51" target="_blank">Amazon Web Services, </a></u><u><a href="https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fionq.com%2Fnews%2Fionq-speeds-quantum-accelerated-drug-development-application-with&esheet=54451094&newsitemid=20260316477473&lan=en-US&anchor=and&index=9&md5=2bed627b388878763f41ffe0e4edb16c" target="_blank">and </a></u><u><a href="https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fionq.com%2Fnews%2Fionq-speeds-quantum-accelerated-drug-development-application-with&esheet=54451094&newsitemid=20260316477473&lan=en-US&anchor=AstraZeneca&index=10&md5=757c1f2c2ec94cf3982700106122164f" target="_blank">AstraZeneca</a></u> achieve 20x performance results and accelerate innovation in drug discovery, materials science, financial modeling, logistics, cybersecurity, and defense. In 2025, the company achieved 99.99% two-qubit gate fidelity, <u><a href="https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fionq.com%2Fnews%2Fionq-achieves-landmark-result-setting-new-world-record-in-quantum-computing&esheet=54451094&newsitemid=20260316477473&lan=en-US&anchor=setting+a+world+record+in+quantum+computing+performance&index=11&md5=32e9cb88afbd439026a352666cc72c91" target="_blank">setting a world record in quantum computing performance</a></u>.<br><br>Headquartered in College Park, Maryland, IonQ has operations in California, Colorado, Massachusetts, Tennessee, Washington, Italy, South Korea, Sweden, Switzerland, Canada, and the United Kingdom. Our quantum computing services are available through all major cloud providers, while we also meet the needs of networking and sensing customers across land, sea, air, and space. IonQ is making quantum platforms more accessible and impactful than ever before. </p></div><p><strong>Location:</strong> Santa Clara, CA<br><strong>Travel:</strong> Up to 25%<strong><br>Job ID:</strong> 1795</p> <p><strong>The Role: <br></strong></p> <p>The Platform Engineering team builds, secures, and operates scalable infrastructure for cloud-managed SaaS products with on-premises components deployed at customer sites. <br><br>The Service Reliability and Operational Intelligence discipline ensures the platform remains stable and resilient, with focus on service continuity and seamless customer experience. It owns production reliability and resilience, observability architecture, service-level objectives, incident response, and implementation of AIOps workflows for triage, remediation, and self-healing. </p> <p>As a Senior Staff Service Reliability and Operational Intelligence Engineer, you define the technical direction for reliability across regions and services. You own the reliability strategy, establish the standards and mechanisms that guide production operations, and elevate excellence through design leadership, operational discipline, and mentorship. You stay deeply hands-on by designing and operating observability platforms, defining and governing SLO programs, leading high-severity incident response, and building resilience and disaster-recovery automation. </p> <p>The work is driven by observability and automation, with a focus on detecting and fixing issues before customers are affected and using every incident to improve the system.</p> <p><strong>Responsibilities</strong>:</p> <ul> <li>Own the technical strategy and multi-year roadmap for operational excellence and production readiness across development, pre-production, and production environments. </li> <li>Define and govern the New Service Introduction framework, including mandatory architecture, security, resilience, capacity, observability, supportability, and release-readiness reviews before services enter production. </li> <li>Establish organization-wide service ownership standards covering service catalog records, accountable owners, dependency maps, runbooks, support models, escalation paths, recovery objectives, and on-call readiness. </li> <li>Lead the architecture and evolution of the shared observability platform, establishing consistent standards for logs, metrics, distributed traces, and profiles across production systems. </li> <li>Define standards for dashboards, alert policies, synthetic monitoring, telemetry quality, retention, sampling, cardinality, and cost controls. </li> <li>Own the reliability governance model for production services, including SLIs, SLOs, error budgets, and escalation mechanisms. </li> <li>Connect service-health signals to customer and business impact, enabling early anomaly detection, service-degradation prevention, and rapid isolation of end-user-impacting events. </li> <li>Advance incident-management maturity through consistent severity classification, incident command, stakeholder and executive communications, automated evidence collection, and coordinated response to high-severity incidents.&