PhysicsX builds physics-based AI for industrial engineering that predicts physical behaviors quickly to accelerate design testing across industries.
<div class="content-intro"><h2>About us</h2> <div>PhysicsX is a deep-tech company with roots in numerical physics and Formula One, dedicated to accelerating hardware innovation at the speed of software.</div> <div>We are building an AI-driven simulation software stack for engineering and manufacturing across advanced industries. By enabling high-fidelity, multi-physics simulation through AI inference across the entire engineering lifecycle, PhysicsX unlocks new levels of optimization and automation in design, manufacturing, and operations — empowering engineers to push the boundaries of possibility. Our customers include leading innovators in Aerospace & Defense, Materials, Energy, Semiconductors, and Automotive.</div></div><h4><strong>Note: </strong>We are currently recruiting for multiple positions, however please only apply for the role that best aligns with your skillset and career goals.</h4> <h1>The Role</h1> <p>The Senior Simulation Data Engineer will extend and operate the infrastructure that powers our research Data Factory. You will be responsible for the end-to-end pipeline: from geometry preparation and simulation orchestration through validation, post-processing, and delivery to downstream ML training systems, using PhysicsX platform orchestration services where synergies exist.</p> <p>This role sits at the intersection of HPC engineering and data engineering. You will orchestrate long-running CFD simulations at scale, build robust data pipelines, and ensure that every simulation we produce meets rigorous quality standards.</p> <h2>Team Context</h2> <p>In this role, you will be vertically embedded in Research , working daily with:</p> <ul> <li><strong>Research Scientists</strong> who define data requirements and quality standards</li> <li><strong>ML Engineers</strong> who consume Data Factory outputs for model training</li> <li><strong>ML Infrastructure Engineers</strong> who are accountable for downstream training infrastructure</li> </ul> <p>You will have end-to-end responsibilities over the Data Factory, with the autonomy to make architectural decisions and the responsibility to keep data flowing reliably.</p> <p>Horizontally, you will be part of an infrastructure engineering group responsible for infrastructure across the company.</p> <h1>What you will do</h1> <h2>Simulation Orchestration</h2> <ul> <li>Extend and operate the Data Factory infrastructure that orchestrates thousands of CFD simulations per day on cloud compute</li> <li>Design and operate job scheduling systems that maximize throughput while handling failures gracefully</li> <li>Build monitoring and alerting to detect simulation failures, convergence issues, and resource bottlenecks early</li> </ul> <h2>Data Pipeline Engineering</h2> <ul> <li>Build high-performance data pipelines that move simulation outputs from solver results to ML-ready training data</li> <li>Implement geometry preprocessing workflows (mesh preparation, morphing, watertightness validation)</li> <li>Design and operate post-processing pipelines: surface decimation, field interpolation, format conversion</li> <li>Optimize I/O performance for large mesh datasets</li> </ul> <h2>Data Quality and Validation</h2> <ul> <li>Implement comprehensive validation checks at every pipeline stage: solver convergence, physical field bounds, post-processing fidelity</li> <li>Build systems that capture and quarantine bad data before they reach training pipelines</li> <li>Track and report data quality metrics across the entire Data Factory</li> <li>Work towards full provenance: training samples should be traceable back to their source geometry and simulation configuration</li> </ul> <h2>Integration and Delivery</h2> <ul> <li>Deliver validated datasets to downstream ML training infrastructure in formats optimized for efficient data loading</li> <li>Design data versioning and cataloging systems that support reproducible training runs</li> <li>Work closely with ML Infrastructure Engineers to ensure smooth handoff between data production and model training</li> <li>Support multi-dataset training workflows</li> </ul> <h1>What you bring to the table</h1> <ul> <li>Ability to scope and effectively deliver projects, prioritising activity as needed.</li> <li>Problem-solving skills and the ability to analyse issues, identify causes, and recommend solutions quickly.</li> <li>Excellent collaboration and communication skills, <strong>especially in a research setting.</strong> You can translate "the model isn't converging" into infrastructure hypotheses and solutions, and can bridge technical abstractions with implementations.</li> <li>5+ years of experience in data engineering, HPC engineering, or simulation infrastructure. <ul> <li>Strong experience with orchestration systems: SLURM, Kubernetes, Temporal</li> <li>Production data pipeline experience: you've built and operated pipelines that process large volumes of data reliably</li> <li>Proficiency in Python for pipeline development and automation</li> <li>Systems engineering fundamentals: Linux, networking, storage systems, performance debugging</li> <li>Experience with cloud infrastructure; ****ideally CoreWeave or similar GPU/HPC-focused clouds</li> <li>Background in HPC for simulation engineering: experience with CFD, FEA, or similar computational workflows (StarCCM+, OpenFOAM, ANSYS, etc.)</li> <li>Experience with geometry processing: mesh manipulation, CAD formats, PyVista</li> <li>Familiarity with scientific data formats: HDF5, VTK, NetCDF, Zarr</li> <li>Data quality engineering experience: validation frameworks, anomaly detection, data observability</li> </ul> </li> </ul> <h2>Ideally</h2> <ul> <li>Understanding of CFD fundamentals, enough to interpret solver outputs and validation metrics</li> <li>Experience with 3D geometry pipelines (mesh decimation, field interpolation)</li> <li>Familiarity with ML data loading patterns and how training systems consume data</li> </ul> <p><strong>What we offer</strong></p