Helsing is a Munich-based defence AI and drone company that develops defence software, drones, and military systems for air, land, sea, and underwater operations.
<h2><strong>Who we are</strong></h2> <div> <p data-renderer-start-pos="859">At Helsing we deliver AI-based capabilities and the enabling foundation that allow machines to perceive and assist human decision-making. You will have the unique opportunity to shape AI capabilities in one of the most challenging sectors, where high generalisation capabilities need to be paired with hardware constraints and robustness against adversarial attacks.</p> <p data-renderer-start-pos="1227"><span class="fabric-text-color-mark" data-renderer-mark="true" data-text-custom-color="#ffc400">You will join a team focused on AI Assurance, where you will develop cutting-edge techniques for scalable evaluation of AI products across the company, design data collection and experimentation strategies to extract causal insights, and enhance responsible decision-making via uncertainty quantification and safety mechanisms.</span></p> </div> <h2 data-renderer-start-pos="4282"><span data-renderer-mark="true" data-mark-type="annotation" data-mark-annotation-type="inlineComment" data-id="e0878c64-297e-4ae8-9721-2e8d280ee236">The Role</span></h2> <p data-renderer-start-pos="99" data-local-id="9eeb02cfa35c">At Helsing we deliver AI-based capabilities and the enabling foundation that allow machines to perceive and assist human decision-making. You will have the unique opportunity to shape AI capabilities in one of the most challenging sectors, where high generalisation capabilities need to be paired with robustness against adversarial attacks and the highest standards of operational safety.</p> <p data-renderer-start-pos="490" data-local-id="a6ed0ca15183">You will be responsible for defining operational domains and evaluating the reliability of AI capabilities developed in-house. Your work will span the full assurance lifecycle: from characterising distribution shifts and failure modes, to developing and extending the state of the art in uncertainty quantification and calibration. You will interface deeply with our AI systems, design rigorous evaluation frameworks, and assess their robustness under real-world and adversarial conditions, collaborating across research, engineering, and product teams to translate assurance findings into actionable improvements.<span id="e0878c64-297e-4ae8-9721-2e8d280ee236" data-renderer-mark="true" data-mark-type="annotation" data-mark-annotation-type="inlineComment" data-id="e0878c64-297e-4ae8-9721-2e8d280ee236"></span></p> <h2 id="You’ll-do-great-if-you….3" data-renderer-start-pos="15247">You should apply if you</h2> <ul class="ak-ul" data-local-id="6905fd6364db" data-indent-level="1"> <li> <p data-renderer-start-pos="1133" data-local-id="32fb9986f92e">Hold an MSc in Mathematics, Statistics, Machine Learning, or a closely related field, with a strong mathematical and statistical foundation.</p> </li> <li> <p data-renderer-start-pos="1277" data-local-id="051e9c28e959">Have hands-on experience in model evaluation, uncertainty quantification, or calibration. You understand the difference between epistemic and aleatoric uncertainty and know how to measure and reduce them in deep learning models.</p> </li> <li> <p data-renderer-start-pos="1509" data-local-id="4aaa93ba608e">Are familiar with methods for distribution shift detection, out-of-distribution detection, and adversarial robustness evaluation, and can design experiments that surface genuine failure modes rather than benchmark artefacts.</p> </li> <li> <p data-renderer-start-pos="1737" data-local-id="36fa69bd64df">Possess solid software engineering skills, writing clean and well-structured code in Python and/or languages like Rust or modern C++, and have experience deploying AI software to production including testing, QA, and monitoring.</p> </li> <li> <p data-renderer-start-pos="1969" data-local-id="e2a7d039b5dd">Have excellent communication skills and the ability to report and present research findings clearly and efficiently, both internally and externally.</p> </li> <li> <p data-renderer-start-pos="2121" data-local-id="55751a445bfc">Are passionate about keeping up to date with current research and enjoy reimplementing and extending state-of-the-art approaches in deep learning evaluation and assurance.</p> </li> </ul> <p><em data-renderer-mark="true"><em>Note: We operate at an intersection where women, as well as other minority groups, are systemically under-represented. We encourage you to apply even if you don’t meet all the listed qualifications; ability and impact cannot be summarised in a few bullet points.</em></em></p> <h2 id="You’ll-be-even-better-prepared-if-you-have-…" data-renderer-start-pos="5667"><span id="b6c557dc-8236-4c01-ad66-49006f89c7cf" data-renderer-mark="true" data-mark-type="annotation" data-mark-annotation-type="inlineComment" data-id="b6c557dc-8236-4c01-ad66-49006f89c7cf">Nice to have</span></h2> <ul class="ak-ul" data-local-id="3d2c8792ff0e" data-indent-level="1"> <li> <p data-renderer-start-pos="2576" data-local-id="90c3546c9563">PhD in model evaluation, uncertainty quantification, robustness, experimental design, causal inference, or a related field, with publications in top-tier venues (e.g. NeurIPS, ICML, ICLR, CVPR).</p> </li> <li> <p data-renderer-start-pos="2774" data-local-id="4ba99d5b0c10">Previous industrial experience assuring the safe deployment of AI products in high-stakes or safety-critical systems.</p> </li> <li> <p data-renderer-start-pos="2895" data-local-id="82b868931094">Familiarity with formal methods, interpretability techniques, or Bayesian approaches to reasoning about model behaviour under uncertainty.</p> </li> <li> <p data-renderer-start-pos="3037" data-local-id="4fd5abc86c87">Experience with adversarial machine learning, red-teaming, or systematic stress-testing of AI systems in operational settings.</p> </li> <li> <p data-renderer-start-pos="3167" data-local-id="b40207366d54">Experience with conformal prediction, calibration methods (e.g. temperature scaling, Platt scaling), or