<div class="content-intro"><p>PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization.</p></div><h2><strong>About the role</strong></h2> <p>PagerDuty’s Operations Cloud runs on a platform that ingests billions of signals and turns them into real-time action for thousands of customers. We’re looking for an early-career AI/ML Engineer who is excited to grow at the intersection of two disciplines: large-scale distributed systems and machine learning.</p> <p>In this role you will help build and ship AI systems that run in production at PagerDuty’s scale — powering Incident Management AI Agents, event intelligence, and the LLM-powered capabilities embedded across our platform. You’ll work alongside senior engineers on real production problems, learning how AI features go from a prototype to something that serves reliably at scale.</p> <p>We are looking for a candidate who is genuinely excited about building with modern AI — LLMs, agents, and retrieval — eager to learn how resilient, high-throughput systems are built, and motivated to grow into an engineer who is strong in both.</p> <h2>What you’ll do</h2> <ul> <li>Contribute to AI-powered features — LLM agents, retrieval, and event intelligence — that operate on high-volume, real-time data, with support and guidance from senior engineers.</li> <li>Help build and maintain the systems behind them — prompt and agent orchestration, retrieval pipelines, tool/API integrations, and inference services — writing code, adding tests, and improving observability.</li> <li>Learn to reason about latency, throughput, cost, and reliability of LLM-powered services, and apply those lessons in the code you write.</li> <li>Help take AI features from prototype toward production, and support the evaluation and monitoring loops that keep them accurate and trustworthy over time.</li> <li>Partner with platform, product, and applied-research teams, asking good questions and turning requirements into working code.</li> <li>Grow through code review, pairing, and mentorship, and steadily take on more ownership as you develop.</li> </ul> <h2>What you’ll bring</h2> <ul> <li>2+ years of software engineering experience building and shipping production software.</li> <li>Degree in CS/a related field, or equivalent practical experience.</li> <li>Solid programming fundamentals and comfort moving between application code and AI/model code.</li> <li>Hands-on experience with modern AI — building something real with LLMs, prompting, retrieval, or agent frameworks.</li> <li>Some exposure to distributed systems and how software runs reliably at scale, and eagerness to deepen it.</li> <li>Strong communication and collaboration skills.</li> </ul> <h2>Nice to have</h2> <ul> <li>Personal, academic, or internship projects involving LLM apps, agents, RAG, or backend services.</li> <li>Exposure to cloud infrastructure (AWS, GCP, or Azure), containers, or Kubernetes.</li> <li>Familiarity with the ecosystem — e.g. LLM APIs and frameworks such as LangChain or LlamaIndex, vector databases, and orchestration/streaming tools like Kafka or Airflow.</li> <li>Interest in agentic systems, evaluation and guardrails for LLMs, or applied problems like anomaly detection and event correlation.</li> <li>Contributions to open-source projects.</li> </ul> <h2>Why PagerDuty</h2> <p>At PagerDuty, AI it’s core to how we help the world’s teams keep their digital services running. This is a place to start your career on problems where scale, latency, and correctness genuinely matter, surrounded by engineers who will invest in helping you grow and who build systems that people depend on in their most critical moments.</p><div class="content-conclusion"><p><strong>Hesitant to apply?</strong></p> <p>We encourage you to submit your resume even if you don't meet every requirement. We value potential and consider each candidate's full professional story. Whether you're exploring a career change or taking your next step, we look forward to reviewing your application. If this just isn’t the right role or time - sign up for <a href="https://careers.pagerduty.com/jobalerts">job alerts</a>!</p> <p><strong>Where we work</strong></p> <p>PagerDuty operates a hybrid work model with <a href="https://careers.pagerduty.com/locations">offices</a> in 8 major cities: Atlanta, Lisbon, London, San Francisco, Santiago, Sydney, Tokyo, and Toronto. While we offer flexibility within our established locations, we <strong>cannot</strong> employ candidates residing in:</p> <p><strong>Location restrictions: </strong><strong><br></strong><strong>Australia:</strong> Northern Territory, Queensland, South Australia, Tasmania, Western Australia<br><strong>Canada:</strong> Alberta, Manitoba, Newfoundland, Northwest Territories, Nunavut, PEI, Quebec, Saskatchewan, Yukon<br><strong>United States:</strong> Alaska, Hawaii, Iowa, Louisiana, Mississippi, Nebraska, New Mexico, Oklahoma, Rhode Island, South Dakota, West Virginia, Wyoming<br><em>Candidates must reside in an eligible location, which vary by role.</em></p> <p><strong>How we work</strong></p> <p><a href="https://careers.pagerduty.com/#values">Our values</a> guide how we support customers