<h3>Company Introduction</h3> <p>At Bot Auto, we are revolutionizing the transportation of goods with our cutting-edge autonomous trucks, enhancing the quality of life for communities around the globe. With the agility of a start-up and the wisdom of seasoned experts, Bot Auto boasts a team that has achieved numerous world-firsts and unparalleled innovations. United by a shared vision, we create miracles and propel the future of transportation. Join us and transform your dreams into reality.</p> <p>You would collaborate with software engineers, AI researchers, and hardware specialists to develop high-performance solutions that meet the stringent requirements of autonomous driving applications. This is an exciting opportunity to work on next-generation transportation technology and make a meaningful impact on the future of mobility.</p> <h3><strong>Key Responsibilities</strong></h3> <ul> <li><strong>Optimize end-to-end GPU performance for real-time autonomous driving workloads</strong>, including sensor processing (e.g., camera, LiDAR) and neural network inference.</li> <li><strong>Develop and optimize parallel computing algorithms</strong> and GPU-accelerated components using technologies such as CUDA.</li> <li>Collaborate with cross-functional teams to <strong>design and improve onboard GPU software architectures</strong> that meet the computational requirements of perception, planning, and control modules.</li> <li><strong>Profile and analyze bottlenecks</strong> across GPU computation, memory access, data movement, synchronization, and CPU–GPU interaction.</li> <li><strong>Debug and optimize GPU-based software</strong> to improve latency, throughput, resource utilization, and runtime stability on embedded platforms.</li> </ul> <h3><strong>Qualifications:</strong></h3> <p><strong>Required</strong>:</p> <ul> <li><strong>Bachelor’s or Master’s degree</strong> in Computer Science, Electrical Engineering, or a related field.</li> <li>Strong knowledge of <strong>parallel computing principles</strong>, GPU architecture, memory hierarchy, and performance optimization techniques.</li> <li>Experience <strong>profiling GPU applications</strong> using tools such as NVIDIA Nsight Systems, Nsight Compute, or equivalent tools.</li> <li>Experience deploying or optimizing neural network <strong>inference workloads</strong> using technologies such as <strong>PyTorch, ONNX, and TensorRT</strong>.</li> <li>Experience with <strong>real-time embedded systems</strong> and handling large data streams from sensors (camera, LiDAR, radar).</li> <li>Strong proficiency in <strong>C/C++ and Python</strong>.</li> </ul> <p><strong>Preferred</strong>:</p> <ul> <li><strong>3+ years of experience</strong> in GPU programming and optimization (e.g., CUDA, OpenCL, Vulkan).</li> <li>Experience with <strong>NVIDIA Jetson Thor, NVIDIA DRIVE Thor</strong>, or similar embedded GPU platforms.</li> <li>Experience with model <strong>quantization</strong>, including FP8 and NVFP4.</li> <li>Experience managing <strong>concurrent GPU workloads</strong> and <strong>resource isolation</strong> using technologies such as NVIDIA Multi-Process Service (<strong>MPS</strong>), Multi-Instance GPU (<strong>MIG</strong>), or other related technologies.</li> <li>Experience with <strong>GPU-accelerated sensor data compression</strong>, including camera, LiDAR, or other onboard sensor data.</li> </ul>