Kodiak Robotics is an autonomous trucking company developing self-driving heavy vehicles and commercial logistics operations.
<div class="content-intro"><p>Kodiak Robotics, Inc. was founded in 2018 and has become a leader in autonomous ground transportation committed to a safer and more efficient future for all. The company has developed an artificial intelligence (AI) powered technology stack purpose-built for commercial trucking and the public sector. The company delivers freight daily for its customers across the southern United States using its autonomous technology. In 2024, Kodiak became the first known company to publicly announce delivering a driverless semi-truck to a customer. Kodiak is also leveraging its commercial self-driving software to develop, test and deploy autonomous capabilities for the U.S. Department of Defense.</p></div><p><span style="font-size: 12pt;">Kodiak's AI is only as good as the speed at which we can train it. Every improvement to our models – from GigaFusionNet to large-scale world models – depends on infrastructure that turns thousands of hours of multimodal driving data into training throughput. We are looking for engineers who make model training fast: streaming massive camera, LiDAR, and radar datasets without stalling a single GPU, sharding data and models efficiently across nodes, and extracting every FLOP from the latest hardware. If you measure your impact in tokens per second and GPU utilization, this role is for you.</span><br><br><span style="font-size: 12pt; font-family: helvetica, arial, sans-serif;"><strong>In this role, you will:</strong></span></p> <ul data-list-tree="true" data-indent="0" data-border="0"> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Design high-throughput data loading and streaming systems for multimodal sensor data (camera, LiDAR, radar), including dataset formats, sharding strategies, and prefetching pipelines that keep GPUs saturated</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Build and optimize distributed training infrastructure across multi-node GPU clusters, applying data, tensor, pipeline, and fully sharded (FSDP/ZeRO) parallelism to models that don't fit on a single device</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Maximize utilization of modern accelerators such as NVIDIA B200s through mixed-precision training (BF16/FP8), fused kernels, memory optimization, and communication/computation overlap</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Profile end-to-end training pipelines to find and eliminate bottlenecks across storage, network, CPU preprocessing, and GPU compute</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Develop scalable dataset construction pipelines that convert petabytes of raw driving logs into training-ready, streamable formats</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Partner with ML teams to scale new architectures from prototype to full-cluster training runs efficiently and reliably</span></li> </ul> <div><span style="font-size: 12pt; font-family: helvetica, arial, sans-serif;"><strong>What you’ll bring:</strong></span></div> <ul data-list-tree="true" data-indent="0" data-border="0"> <li style="font-size: 12pt;"><span style="font-size: 12pt;">BS, MS, or PhD in Computer Science or a related field, and at least 2-3 years of industry experience in ML systems or infrastructure</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Hands-on experience with distributed training frameworks and techniques (PyTorch DDP/FSDP, DeepSpeed, Megatron, NCCL) and a strong grasp of parallelism trade-offs</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Experience building high-performance data pipelines for large-scale training, including streaming dataset formats (WebDataset, MosaicML Streaming/MDS, or similar), sharding, and storage/network-aware loading</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Deep understanding of GPU performance: mixed precision, memory hierarchy, kernel fusion, profiling tools (Nsight, PyTorch Profiler), and interconnects (NVLink, InfiniBand)</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Strong Python skills and proficiency in PyTorch internals; systems-level experience (C++/CUDA/Triton) a plus</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Passion for building the infrastructure that lets AI for the physical world train faster, scale further, and improve continuously</span></li> </ul> <p class="p2"><span style="font-size: 12pt; font-family: helvetica, arial, sans-serif;"><strong>What we offer:</strong></span></p> <ul> <li style="font-size: 12pt; font-family: helvetica, arial, sans-serif;"><span style="font-size: 12pt; font-family: helvetica, arial, sans-serif;">Competitive compensation package including equity and annual bonuses</span></li> <li style="font-size: 12pt; font-family: helvetica, arial, sans-serif;"><span style="font-size: 12pt; font-family: helvetica, arial, sans-serif;">Excellent Medical, Dental, and Vision plans through Kaiser Permanente, Cigna, and MetLife (including a medical plan with infertility benefits)</span></li> <li style="font-size: 12pt; font-family: helvetica, arial, sans-serif;"><span style="font-size: 12pt; font-family: helvetica, arial, sans-serif;">MetLife Legal Services, Identity & Fraud Protection, Hospital Indemnity Insurance, Accident Insurance, & Critical Illness Insurance</span></li> <li style="font-size: 12pt; font-family: helvetica, arial, sans-serif;"><span style="font-size: 12pt; font-family: helvetica, arial, sans-serif;">Flexible PTO, 10 paid holidays, and generous parental leave policies</span></li> <li style="font-size: 12pt; font-family: helvetica, arial, sans-serif;"><span style="font-size: 12pt; font-family: helvetica, arial, sans-serif;">Our office is centrally located in Mountain View, CA</span></li> <li style="font-size: 12pt; font-family: helvetica, arial, sans-serif;"><span