Graphcore is an AI infrastructure and hardware company.
<h1>About us</h1> <p>We are looking for a disciplined and dynamic Systems Engineer with focus on server CPU based system to join our growing compute rack validation team. Candidate we are seeking should have demonstrated work-experience in leading server rack and blade hardware systems deployment, hardware installation, and inventory management activities in the Austin, TX area. As a diligent leader in Systems Engineering, you will drive multiple aspects of post-silicon validation throughout the life cycle of the program. In this high visibility position, you will be part of a technical team chartered to innovate and improve system bring-up and enablement capabilities, as well as silicon and system validation to deliver the highest quality, industry leading technologies to market. Your technical leadership skills, systems engineering and hardware bring-up, validation and debug expertise will be necessary towards product development, definition, root cause and resolution. Your agility and collaborative approach will be essential to work within System Validation & other engineering teams (System Architects, SoC and Rack FW etc).</p> <p>The ideal candidate will be driving key areas around at-scale system validation including ARM based server and rack level systems bring-up (nodes and rack level systems). Candidate will be immersed in challenging system enablement work, ramp-up post-silicon capabilities in engineering lab environments, validation tests execution/triage. The candidate will be leading contributor towards state-of-the-art HW bring-up and lab capabilities for Grapchore’s system engineering. The candidate should be able to work in a global environment while maintaining a synergetic culture.</p> <p> </p> <p><strong>Primary Responsibilities:</strong></p> <ul> <li>Install, configure, commission (and decommission if needed) blade servers, chassis, switches, and supporting infrastructure.</li> <li>Lead rack and stack activities, including mounting equipment, cable management, and labelling. Execute hardware upgrades, replacements, and troubleshooting of server and network components.</li> <li>Maintain accurate asset records within DCIM platforms and inventory management systems.</li> <li>Conduct physical audits and reconcile inventory discrepancies.</li> <li>Track hardware movements, deployments, and decommissions through established change management processes.</li> <li>Document installation procedures, rack layouts, cabling diagrams, and inventory updates.</li> <li>Support data center migration, expansion, and refresh projects.</li> <li>Collaborate with engineering, operations, logistics, and project management teams.</li> <li>Adhere to all data center safety, security, and operational standards.</li> <li>Develop, setup and scale key methodologies for at-scale test execution, lab HW and system SW capabilities as well as system visibilities and debug tools necessary for successful system (HW/SW/FW) bring-up and system validation at blade and rack level for AI compute rack.</li> <li>Ability to work independently in a production ready environment, and a commitment tomaintaining accurate inventory and asset records.</li> <li>Triage issues found during server rack validation bring-up, Post-Silicon Validation, and production phases of the program. Ensure issues are solved on time with quality.</li> <li>Lead test execution of key domains within AI compute solutions like CPU, GPU, memory, HBM, IO etc.</li> <li>Drive technical innovation to improve capabilities across system validation, including tools, script development, technical and procedural methodology enhancement, and various internal and cross-functional technical initiatives.</li> </ul> <p><strong>Qualifications:</strong></p> <ul> <li>Strong analytical/problem-solving skills and pronounced attention to details</li> <li>Experience in Blade server installation and maintenance (Cisco UCS, HPE Synergy, Dell MX, or similar).</li> <li>Rack and stack deployments in enterprise or hyperscale environments.</li> <li>Copper and fiber cabling installation and management.</li> <li>DCIM and asset management platforms.</li> <li>Strong understanding of server, storage, and networking hardware.</li> <li>Experience performing inventory audits and maintaining asset accuracy.</li> <li>Ability to read rack elevation diagrams, cabling schematics, and deployment documentation.</li> <li>Familiarity with ticketing and change management systems.</li> <li>Exposure to Linux (ubuntu) OS bootable images and system firmware basics for image building, provisioning and firmware flashing.</li> <li>Exposure to automation testing, to enable execution of hardware acceptance tests, best-known-config testing etc.</li> <li>Exposure to python script development and execution.</li> <li>Proven experience in understanding, defining and enabling storage (storage rack), networking capabilities (network rack, DNS, DHCP etc) in a lab environment to help add end-to-end validation and debug capabilities for rack and blade validation.</li> <li>Excellent communication and coordination skills.</li> <li>Detailed oriented, highly organized, able to prioritize, and juggle multiple work streams to tight deadlines.</li> <li>Technical leadership: capable of championing new tools, methods, and capabilities to drive platform validation improvements in schedule, quality, or coverage.</li> <li>Experience working with data center technical staff, 3<sup>rd</sup> party vendors, ODMs etc throughout the life cycle of server system product development.</li> <li>Must be a self-starter, and able to independently drive tasks to completion</li> </ul> <p><strong>Preferred Qualifications:</strong></p> <ul> <li>Masters or PhD in Electrical Engineering, Computer Engineering or a related field.</li> <li>10+ years of work experience demonstrating working on complex systems engineering challenges to validate and debug HW-FW-SW challenges in a server compute rack or data center blade