Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
at Didi · 501-1000 employees
- Seniority
- Staff Principal
- Location
- San Jose, CA
- Posted
- 5d ago
at Didi · 501-1000 employees
Didi Autonomous Driving develops Level-4 autonomous driving technology and unmanned ride-hailing vehicles for urban mobility.
<p class="ql-direction-ltr ql-long-10000202573"><strong class="ql-author-10000202573 ql-size-12 ql-font-arial">About the Company</strong></p> <p class="ql-direction-ltr ql-long-10000202573"><span class="ql-author-10000202573 ql-size-12 ql-font-arial">DiDi's autonomous driving unit was established in 2016 with the mission of developing </span><span class="ql-author-10000202573 ql-size-12 ql-font-arial">Level 4 autonomous driving (AD) technology to make transportation safer and more efficient. In August 2019, the unit became an independent company, DiDi Autonomous Driving, dedicated to advanced AD R&D, product application, and business expansion. </span><span class="ql-size-12 ql-font-microsoftyahei ql-author-10000202573">We believe integrating AD technology into a shared-mobility fleet will generate immense social value. By leveraging DiDi's specialized technology, operational expertise, and integrated ecosystem, we are positioned to build and operate a highly efficient, user-oriented autonomous fleet.</span></p> <div class="ql-direction-ltr ql-long-10000202573" data-header="4" data-foldable="true" data-default-linespacing="100"> <p data-path-to-node="1"> </p> <p data-path-to-node="1"><strong data-path-to-node="1" data-index-in-node="0">About The Role</strong></p> <p data-path-to-node="2">We are seeking an experienced and mission-driven Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization to lead the performance tuning, deployment, and resource scheduling of cutting-edge AI models across on-vehicle and cloud infrastructure. In this role, you will design high-efficiency inference pipelines, build system-level stability frameworks, and optimize hardware execution to ensure ultra-low latency and rock-solid operational reliability. You will act as a technical leader in AI infrastructure, accelerating model iteration and bridging the gap between frontier deep learning algorithms and real-time autonomous systems.</p> <p data-path-to-node="3"> </p> <p data-path-to-node="3"><strong data-path-to-node="3" data-index-in-node="0">Responsibilities</strong></p> <ul data-path-to-node="4"> <li> <p data-path-to-node="4,0,0">Own the deployment, optimization, and resource scheduling of vehicle-side AI models, ensuring high efficiency, low latency, and robust execution within embedded constraints.</p> </li> <li> <p data-path-to-node="4,1,0">Lead vehicle-side system stability initiatives, conducting independent root-cause analysis and driving resolution for complex, system-level performance bottlenecks and runtime anomalies.</p> </li> <li> <p data-path-to-node="4,2,0">Architect and scale service-oriented deployment environments for Large Language Models (LLMs) and foundational models to support offline simulation, automated annotation, and rapid model validation.</p> </li> <li> <p data-path-to-node="4,3,0">Track and evaluate cutting-edge industry methodologies, continuously integrating advanced optimization toolchains, quantization techniques, and execution engines.</p> </li> <li> <p data-path-to-node="4,4,0">Establish system-level profiling and telemetry frameworks using CUDA tools to monitor, analyze, and maximize hardware utilization across target GPU architectures.</p> </li> <li> <p data-path-to-node="4,5,0">Collaborate cross-functionally with Autonomous Driving Perception/Prediction, Cloud Infrastructure, and Safety teams to enable rapid algorithm iteration and scalable vehicle deployment.</p> </li> </ul> <p data-path-to-node="5"> </p> <p data-path-to-node="5"><strong data-path-to-node="5" data-index-in-node="0">Qualifications</strong></p> <ul data-path-to-node="6"> <li> <p data-path-to-node="6,0,0">Master’s or higher degree in Computer Science, Software Engineering, Systems Engineering, or a closely related technical field.</p> </li> <li> <p data-path-to-node="6,1,0">3-8+ years of industry experience in high-performance computing, AI infrastructure, model optimization, or embedded deployment.</p> </li> <li> <p data-path-to-node="6,2,0">Strong proficiency in C++ and Python, with solid expertise in parallel programming (CUDA, OpenMP) and low-level system profiling tools.</p> </li> <li> <p data-path-to-node="6,3,0">Deep familiarity with mainstream inference engines (e.g., TensorRT, ONNX Runtime) and specialized LLM inference/serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM).</p> </li> <li> <p data-path-to-node="6,4,0">Practical understanding of modern GPU hardware architectures (e.g., NVIDIA Hopper, Thor) and memory bandwidth management.</p> </li> <li> <p data-path-to-node="6,5,0">Demonstrated ability to diagnose complex software-hardware integration issues and drive scalable, production-grade solutions.</p> </li> </ul> <p data-path-to-node="7"><strong data-path-to-node="7" data-index-in-node="0">Preferred Qualifications</strong></p> <ul data-path-to-node="8"> <li> <p data-path-to-node="8,0,0">Hands-on experience optimizing and deploying AI models on the NVIDIA Thor platform, including hardware resource scheduling and acceleration.</p> </li> <li> <p data-path-to-node="8,1,0">Proven track record of serving large foundation models (e.g., LLaMA, Qwen, GPT) in production or high-throughput cloud pipelines using frameworks like vLLM, SGLang, TGI, or LightLLM.</p> </li> <li> <p data-path-to-node="8,2,0">Background in deep learning training frameworks (PyTorch) and practical experience with model quantization (INT8/FP8/AWQ), kernel fusion, or graph compilation.</p> </li> <li> <p data-path-to-node="8,3,0">Experience deploying real-time, high-availability AI workloads in autonomous vehicles, robotics, or edge devices.</p> </li> </ul> </div> <p class="ql-direction-ltr ql-long-10000202573"> </p> <p class="ql-direction-ltr ql-long-10000202573"><span class="ql-author-10000202573 ql-size-12 ql-font-arial">The base salary range for this full-time position is $169,783 - $351,000 annually in addition to bonus, equity and benefits. Our salary ranges are determined by role, le