Senior Software Engineer
at Sonyinteractiveentertainmentglobal
- Seniority
- Senior
- Location
- United States, San Mateo, CA
- Posted
- 6d ago
at Sonyinteractiveentertainmentglobal
<div class="content-intro"><p><strong>Why Sony Interactive Entertainment?</strong></p> <p>Sony Interactive Entertainment isn’t just the Best Place to Play — it’s also the Best Place to Work. Sony Interactive Entertainment (SIE) is the company behind the PlayStation brand. As a subsidiary of Sony Group Corporation, we’re part of a proud legacy of innovation and excellence. SIE is a dynamic technology company, delivering cutting-edge hardware and network services to more than 100 million people and an entertainment leader, home to some of the most beloved and recognizable intellectual properties (IP) in the world. Our role at SIE is to create and nurture the experiences under the PlayStation brand, a name synonymous with entertainment excellence and creativity.</p></div><h1><span style="color: #0f4761;">Senior Software Engineer – Logging, Diagnostics & Post-Release Incident Analysis </span></h1> <p style="color: !important;"><span style="color: #0f4761;">Overview </span></p> <p style="color: !important;"><span style="color: #000000;">The PlayStation Portal Platform team is seeking a Senior Software Engineer to lead the design, development, and operation of Logging and Diagnostics infrastructure that enables rapid detection, investigation, and resolution of Performance, Stability, and Platform Quality issues occurring after products are released to customers. </span></p> <p style="color: !important;"><span style="color: #000000;">In this role, you will be responsible for building scalable Logging Architecture, Crash Diagnostics, and Incident Analysis Frameworks that empower engineering teams to perform efficient root cause analysis of field issues. You will collaborate closely with Engineering, Quality Engineering (QE), Customer Support, Security, and Product teams to reduce the time from issue detection to root cause identification and drive continuous improvements in product quality and reliability. </span></p> <p style="color: !important;"><span style="color: #000000;">We are looking for engineers who are passionate about quality, thrive on solving complex system-level problems, and take strong ownership in delivering exceptional user experiences through continuous improvement. </span></p> <p style="color: !important;"><span style="color: #0f4761;">Key Responsibilities </span></p> <p style="color: !important;"><span style="color: #0f4761;"><strong>Logging Architecture</strong> </span></p> <ul> <li><span style="color: #000000;">Define and drive the overall platform logging strategy. </span></li> </ul> <ul> <li><span style="color: #000000;">Design and implement logging frameworks for AOSP-based systems. </span></li> </ul> <ul> <li><span style="color: #000000;">Establish log collection, storage, retention, and archival strategies. </span></li> </ul> <ul> <li><span style="color: #000000;">Define structured logging standards and best practices. </span></li> </ul> <ul> <li><span style="color: #000000;">Design logging mechanisms optimized for field debugging and post-release analysis. </span></li> </ul> <ul> <li><span style="color: #000000;">Improve platform observability and telemetry quality. </span></li> </ul> <p style="color: !important;"><span style="color: #0f4761;"><strong>Diagnostics & Debuggability</strong> </span></p> <ul> <li><span style="color: #000000;">Build infrastructure for Crash, ANR, Native Crash, and Kernel Panic analysis. </span></li> </ul> <ul> <li><span style="color: #000000;">Define strategies for collecting and analyzing Tombstones, DropBox entries, Bugreports, and diagnostic artifacts. </span></li> </ul> <ul> <li><span style="color: #000000;">Develop diagnostic frameworks for system failure analysis and troubleshooting. </span></li> </ul> <ul> <li><span style="color: #000000;">Create tools that support issue reproduction and root cause analysis. </span></li> </ul> <ul> <li><span style="color: #000000;">Drive initiatives that improve platform debuggability and diagnosability. </span></li> </ul> <ul> <li><span style="color: #000000;">Partner with development teams to ensure diagnostic capabilities are built into the platform. </span></li> </ul> <p style="color: !important;"><span style="color: #0f4761;"><strong>Post-Release Incident Analysis</strong> </span></p> <ul> <li><span style="color: #000000;">Investigate and analyze production issues occurring in customer environments. </span></li> </ul> <ul> <li><span style="color: #000000;">Identify, classify, and track recurring failure patterns. </span></li> </ul> <ul> <li><span style="color: #000000;">Lead incident reviews, postmortems, and corrective action initiatives. </span></li> </ul> <ul> <li><span style="color: #000000;">Drive improvements in Mean Time to Detect (MTTD) and Mean Time to Resolution (MTTR). </span></li> </ul> <ul> <li><span style="color: #000000;">Provide actionable recommendations to engineering teams for quality improvements. </span></li> </ul> <p style="color: !important;"><span style="color: #0f4761;"><strong>Monitoring & Alerting</strong> </span></p> <ul> <li><span style="color: #000000;">Design and implement log-based monitoring systems. </span></li> </ul> <ul> <li><span style="color: #000000;">Develop anomaly detection mechanisms for platform health and stability. </span></li> </ul> <ul> <li><span style="color: #000000;">Build automated alerting and escalation frameworks. </span></li> </ul> <ul> <li><span style="color: #000000;">Create dashboards and reporting tools for field quality monitoring. </span></li> </ul> <ul> <li><span style="color: #000000;">Define and track reliability metrics and operational KPIs. </span></li> </ul> <p style="color: !important;"><span style="color: #0f4761;"><strong>Security & Compliance</strong> </span></p> <ul> <li><span style="color: #000000;">Design and implement secure logging and diag