RadixArk develops open-source and commercial tools that accelerate the inference and continual optimization of large AI models.
<div data-page-id="R5pEdqmJmoX3WtxIJiNjDTCKpUb" data-lark-html-role="root" data-docx-has-block-data="false"> <h3 class="heading-3 ace-line old-record-id-IdxIdWSMyoqD4TxNEYpjwCiDpQe">About the Role</h3> <div class="ace-line ace-line old-record-id-RCDod239Zo8KBZxH3wZjB3TUp1f">RadixArk is looking for a Member of Technical Staff — Backend/API Platform Engineer to build the API layer, control plane, and platform services that power SGLang and Miles in production. You'll design and implement the REST/gRPC APIs, authentication systems, multi-tenancy isolation, and monitoring infrastructure that thousands of developers and companies rely on. This role bridges high-performance inference/training systems with production-grade platform engineering.</div> <h3 class="heading-3 ace-line old-record-id-PNQ0dhVtqouI29xSCxPjofzMpje">Requirements</h3> <ul class="list-bullet1"> <li class="ace-line ace-line old-record-id-Eiihd6nU7oUX3UxOg3cjeMESpld" data-list="bullet"> <div>4+ years experience building production backend systems, APIs, or platform infrastructure</div> </li> <li class="ace-line ace-line old-record-id-XYQQdWbb6oOmz9xzJtMjLLB1pJe" data-list="bullet"> <div>Bachelor's or Master's degree in Computer Science, Engineering, or equivalent industry experience</div> </li> <li class="ace-line ace-line old-record-id-EmRSdMYDWoQWP1xbYdMj8wFhppb" data-list="bullet"> <div>Strong proficiency in Python, Go, or Rust with production-quality code standards</div> </li> <li class="ace-line ace-line old-record-id-MZKXdfn9No00PcxiXMOjbbYpp3c" data-list="bullet"> <div>Experience designing and building REST/gRPC APIs at scale</div> </li> <li class="ace-line ace-line old-record-id-SGvadeTUTodQwTxmbMajSKQRpWb" data-list="bullet"> <div>Solid understanding of distributed systems, databases, caching, and message queues</div> </li> <li class="ace-line ace-line old-record-id-M2uAd86XDogv4LxbfR9js1iWp8e" data-list="bullet"> <div>Experience with authentication, authorization, rate limiting, and multi-tenancy</div> </li> <li class="ace-line ace-line old-record-id-LSJ7d2xreoQn7yxTM1ajdvqqpYL" data-list="bullet"> <div>Familiarity with cloud platforms (AWS, GCP, Azure) and Kubernetes</div> </li> <li class="ace-line ace-line old-record-id-GXImd4spnooU6Yxs4l5jRG8Upsb" data-list="bullet"> <div>Experience with monitoring and observability tools (Prometheus, Grafana, DataDog)</div> </li> <li class="ace-line ace-line old-record-id-RpkVdjqEJoGFMLxZGoBjRvblped" data-list="bullet"> <div>Understanding of ML serving infrastructure or high-throughput systems is a plus</div> </li> </ul> <h3 class="heading-3 ace-line old-record-id-QRJ6dESr4o0h2WxRCBtjdgQ4pqb">Responsibilities</h3> <ul class="list-bullet1"> <li class="ace-line ace-line old-record-id-VZ6tdmktcoSXVsx9b58jariqpFz" data-list="bullet"> <div>Design and build production APIs for SGLang and Miles: REST/gRPC endpoints, client SDKs, API versioning</div> </li> <li class="ace-line ace-line old-record-id-AiqodVosOoOCrKxZh4YjXTtApHh" data-list="bullet"> <div>Implement authentication, authorization, and rate limiting systems for multi-tenant deployments</div> </li> <li class="ace-line ace-line old-record-id-IyR7drVF9oorY7xdW6tjC1DJpag" data-list="bullet"> <div>Build control plane infrastructure: job scheduling, resource allocation, model deployment management</div> </li> <li class="ace-line ace-line old-record-id-HDSsd8WBhoeh8ixTZ9WjC2e9pie" data-list="bullet"> <div>Create monitoring, logging, and observability systems for production inference and training workloads</div> </li> <li class="ace-line ace-line old-record-id-Eqb1do0Gwod4RixiCEqjyKs1pZe" data-list="bullet"> <div>Design and implement billing integration, usage tracking, and quota management</div> </li> <li class="ace-line ace-line old-record-id-EhAzd1K1ooC540xtksPjVVXlpTb" data-list="bullet"> <div>Build management dashboards and admin tools for cluster operations</div> </li> <li class="ace-line ace-line old-record-id-UoKcdj6apoxlfNxAPF5jf9qgpnf" data-list="bullet"> <div>Ensure API reliability, performance, and security at scale</div> </li> <li class="ace-line ace-line old-record-id-RGqcdf8eeoBHGixra81jqZPIpIb" data-list="bullet"> <div>Implement multi-tenancy isolation and security boundaries</div> </li> <li class="ace-line ace-line old-record-id-CU25dCrdfojcPlxYRv8jl64Lpic" data-list="bullet"> <div>Create deployment automation, CI/CD pipelines, and rollback procedures</div> </li> <li class="ace-line ace-line old-record-id-PfQwdhSQnosCH2xZxfgjtKCFpjc" data-list="bullet"> <div>Write comprehensive API documentation and integration guides</div> </li> <li class="ace-line ace-line old-record-id-VINRdSK7IokjVlx7yOSjp32rp4g" data-list="bullet"> <div>Partner with Systems Engineers to optimize end-to-end latency from API → serving layer</div> </li> <li class="ace-line ace-line old-record-id-GHgpdRDW5oFCMBxoDc0jP4rjpGe" data-list="bullet"> <div>Debug production issues and implement reliability improvements <div data-page-id="R5pEdqmJmoX3WtxIJiNjDTCKpUb" data-lark-html-role="root" data-docx-has-block-data="false"> <h3>About RadixArk</h3> <p>RadixArk is an infrastructure-first company built by engineers who've shipped production AI systems, created SGLang (30K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). Founded by AI infrastructure veterans from xAI and NVIDIA, we're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs.</p> <h3>Compensation</h3> <p>We offer competitive compensation with equity, comprehensive health benefits, and flexible work arrangements. Compensation is determined by location, level, and experience.</p> <h3>Equal Opportunity</h3> <p>RadixArk is an Equal Opportunity Employer and is proud