Kodiak Robotics is an autonomous trucking company developing self-driving heavy vehicles and commercial logistics operations.
<div class="content-intro"><p>Kodiak Robotics, Inc. was founded in 2018 and has become a leader in autonomous ground transportation committed to a safer and more efficient future for all. The company has developed an artificial intelligence (AI) powered technology stack purpose-built for commercial trucking and the public sector. The company delivers freight daily for its customers across the southern United States using its autonomous technology. In 2024, Kodiak became the first known company to publicly announce delivering a driverless semi-truck to a customer. Kodiak is also leveraging its commercial self-driving software to develop, test and deploy autonomous capabilities for the U.S. Department of Defense.</p></div><div><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Kodiak is building AI that doesn't just perceive the world, it learns how the physics of the world works. We are developing large-scale generative world models that learn to predict realistic, physically consistent futures from real-world sensor data. This capability serves as the foundation for scalable closed-loop training, validation, and long-tail scenario generation, and is distilled into the onboard models that drive our autonomous trucks. We are looking for a research scientist to lead the design and development of world models capable of generating multi-sensor, multi-view, temporally coherent driving scenarios conditioned on actions, 3D scene context, and text.</span></div> <div> </div> <div><strong><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">In this role, you will:</span></strong></div> <ul> <li style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Design and train generative world models that synthesize realistic multi-camera video and LiDAR conditioned on ego trajectories, 3D scene context, and text</span></li> <li style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Research and implement conditional diffusion architectures for driving, including spatiotemporal attention, latent space design, and action-conditioned generation</span></li> <li style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Develop techniques for multi-view geometric consistency in generated outputs, drawing on neural rendering, cross-view attention, and 3D-aware generative approaches</span></li> <li style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Build methods for joint multimodal generation that maintain cross-sensor consistency between camera, LiDAR, and radar outputs</span></li> <li style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Design evaluation frameworks that measure world model quality beyond pixel-level metrics, including scenario fidelity and autoregressive stability</span></li> <li style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Scale training pipelines to learn from thousands of hours of real-world driving data across multiple sensor modalities</span></li> </ul> <div><strong><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">What you'll bring:</span></strong></div> <ul> <li style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">PhD in Computer Science, AI, Robotics, or a related field, with a focus on generative modeling, neural rendering, or video synthesis</span></li> <li style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Strong publication record or demonstrated research contributions in diffusion models, video generation, neural radiance fields, 3D-aware generative models, or world models</span></li> <li style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Experience with neural rendering and view synthesis and an understanding of multi-view geometric consistency</span></li> <li style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Proficiency working with multimodal sensor data (camera, LiDAR, radar) and familiarity with 3D representations such as BEV grids, voxel fields, or tri-planes</span></li> <li style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Strong implementation skills in Python and PyTorch, with experience training large generative models at scale using distributed training</span></li> <li style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Passion for building AI that understands and predicts the physical world to enable safe autonomous driving</span></li> </ul> <p><strong><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">What We Offer:</span></strong></p> <ul> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Competitive compensation package including equity and annual bonuses</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Excellent Medical, Dental, and Vision plans through Kaiser Permanente, Cigna, and MetLife (including a medical plan with infertility benefits)</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">MetLife Legal Services, Identity & Fraud Protection, Hospita