RadixArk develops open-source and commercial tools that accelerate the inference and continual optimization of large AI models.
<h2 data-start="201" data-end="222"><strong data-start="204" data-end="222">About the Role</strong></h2> <p data-start="224" data-end="405">RadixArk is looking for a<span class="Apple-converted-space"> </span>Member of Technical Staff Cluster Infrastructure<span class="Apple-converted-space"> </span>to architect and scale the core compute platform that powers frontier-level AI training and inference.</p> <p data-start="407" data-end="634">You will design and operate highly reliable, high-performance GPU/TPU clusters, build next-generation scheduling and resource management systems, and push the limits of large-scale distributed infrastructure for AI workloads.</p> <p data-start="636" data-end="854">This role focuses on deep systems engineering across cluster architecture, networking, scheduling, and performance optimization. Your work will directly impact how efficiently frontier AI models are trained and served.</p> <h2 data-start="861" data-end="880"><strong data-start="864" data-end="880">Requirements</strong></h2> <ul data-start="882" data-end="1536"> <li data-start="882" data-end="981"> <p data-start="884" data-end="981">5+ years of experience in distributed systems, infrastructure, or large-scale compute platforms</p> </li> <li data-start="982" data-end="1058"> <p data-start="984" data-end="1058">Strong background in distributed systems design and systems architecture</p> </li> <li data-start="1059" data-end="1157"> <p data-start="1061" data-end="1157">Deep experience with cluster management systems (Kubernetes, Slurm, Ray, or custom schedulers)</p> </li> <li data-start="1158" data-end="1236"> <p data-start="1160" data-end="1236">Hands-on experience with GPU/TPU infrastructure in production environments</p> </li> <li data-start="1237" data-end="1289"> <p data-start="1239" data-end="1289">Strong Linux systems and networking fundamentals</p> </li> <li data-start="1290" data-end="1356"> <p data-start="1292" data-end="1356">Proficiency in Go, Rust, C++, or Python for production systems</p> </li> <li data-start="1357" data-end="1466"> <p data-start="1359" data-end="1466">Experience debugging complex multi-layer issues across hardware, OS, networking, and distributed services</p> </li> <li data-start="1467" data-end="1536"> <p data-start="1469" data-end="1536">Proven ability to design reliable, scalable systems in production</p> </li> </ul> <p data-start="1538" data-end="1554"><strong data-start="1538" data-end="1554">Strong Plus:</strong></p> <ul data-start="1556" data-end="1837"> <li data-start="1556" data-end="1603"> <p data-start="1558" data-end="1603">Experience with large-scale ML/AI workloads</p> </li> <li data-start="1604" data-end="1673"> <p data-start="1606" data-end="1673">Familiarity with RDMA, InfiniBand, or high-performance networking</p> </li> <li data-start="1674" data-end="1726"> <p data-start="1676" data-end="1726">Experience operating clusters at 1000+ GPU scale</p> </li> <li data-start="1727" data-end="1780"> <p data-start="1729" data-end="1780">Background in HPC or performance-critical systems</p> </li> <li data-start="1781" data-end="1837"> <p data-start="1783" data-end="1837">Open-source contributions in systems or infrastructure</p> </li> </ul> <h2 data-start="1844" data-end="1867"><strong data-start="1847" data-end="1867">Responsibilities</strong></h2> <ul data-start="1869" data-end="2606"> <li data-start="1869" data-end="1945"> <p data-start="1871" data-end="1945">Architect and scale large AI compute clusters for training and inference</p> </li> <li data-start="1946" data-end="2020"> <p data-start="1948" data-end="2020">Design cluster management, scheduling, and resource allocation systems</p> </li> <li data-start="2021" data-end="2095"> <p data-start="2023" data-end="2095">Optimize performance, utilization, and reliability of GPU/TPU clusters</p> </li> <li data-start="2096" data-end="2154"> <p data-start="2098" data-end="2154">Improve fault tolerance and system resilience at scale</p> </li> <li data-start="2155" data-end="2244"> <p data-start="2157" data-end="2244">Drive observability, monitoring, and performance profiling for cluster infrastructure</p> </li> <li data-start="2245" data-end="2323"> <p data-start="2247" data-end="2323">Collaborate with ML and systems engineers to support frontier AI workloads</p> </li> <li data-start="2324" data-end="2388"> <p data-start="2326" data-end="2388">Lead capacity planning and infrastructure scaling strategies</p> </li> <li data-start="2389" data-end="2463"> <p data-start="2391" data-end="2463">Build internal platforms and tooling to improve developer productivity</p> </li> <li data-start="2464" data-end="2540"> <p data-start="2466" data-end="2540">Document architecture, operational practices, and reliability strategies</p> </li> <li data-start="2541" data-end="2606"> <p data-start="2543" data-end="2606">Contribute to long-term platform vision and technical direction</p> </li> </ul> <h3> </h3> <h3>About RadixArk</h3> <p>RadixArk is an infrastructure-first company built by engineers who've shipped production AI systems, created SGLang (30K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). Founded by AI infrastructure veterans from xAI and NVIDIA, we're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs.</p> <h3>Compensation</h3> <p>We offer competitive compensation with equity, comprehensive health benefits, and flexible work arrangements. Compensation is determined by location, level, and experience.</p> <h3>Equal Opportunity</h3> <p>RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, anc