H1 is a healthcare data platform that sells detailed information about doctors to pharmaceutical companies, hospital systems, and health insurers.
H1's mission is to connect the world to the right doctor. We've built one of the largest healthcare datasets in the world: profiles of more than 10 million physicians, assembled from billions of insurance claims, 25 million publications, and nearly 500K clinical trials, with hundreds of sources feeding each doctor's profile. 85% of the top 20 pharma companies use it to decide who should run their clinical trials. Nine of the ten largest US health insurers use it to build their networks. And increasingly, the leading frontier models use our data to help patients find the right doctor. Most people we help never see our name. The engineering behind this is genuinely hard: thousands of sources, no shared identifiers, strict privacy rules, and answers that have to be right. Data Engineering builds that dataset. The team owns the pipelines that turn raw sources into the physician profiles our customers query: hundreds of terabytes on PySpark, EMR, and Hudi. The day-to-day is deciding how the truth gets computed: which source wins when two disagree, how the same doctor gets matched across systems that have never heard of each other, and how it all stays fresh and affordable at scale. Our data is growing faster than our team, so we're hiring engineers to own the hardest pipelines end to end. WHAT YOU'LL DO AT H1 You'll be the senior-most data engineer on the Real World Evidence (RWE) team, owning the claims pipelines that are becoming H1's largest and most important datasets. This is a hands-on role: you'll be in the codebase every day. What's claims data? Nearly every doctor visit in America generates an insurance claim: who was treated, for what, by whom. Billions of them, from thousands of sources that share no common format or identifier. Assembled correctly, they reveal how medicine is actually practiced, which no other dataset can show. You will: - Claims pipelines: turn billions of US insurance claims into the physician-level data our customers query every day. This is the largest dataset H1 has ever built, and it's yours. - Patient journeys: reconstruct treatment timelines from claims and health records, so customers can see how patients actually move through care and which doctors to reach. - Performance and cost: own the hundreds-of-terabytes Spark workloads end to end, and make them faster, more reliable, and cheaper. - Roadmap: set the technical direction for RWE data with Product, Data Science, and downstream teams, and make the architecture calls. - Mentorship: raise the team's bar through design reviews, pairing, and deep domain teaching. ABOUT YOU You've run data products at a serious scale, and you liked owning all of it: the design, implementation, the tradeoffs, the incident at 2am, the cost curve. You're at your best when the problem is ambiguous and the data is messy, and you'd rather ship a pragmatic call this quarter than a perfect one next year. You've probably never worked in healthcare. That's fine. The engineers who do this well came for the problem. REQUIREMENTS - 8+ years building production data or backend systems, and you still write code daily and want to keep it that way. - You've designed and run Spark pipelines at scale and owned the performance, cost, and reliability tradeoffs yourself. - Leadership: you've driven multi-quarter, cross-team technical work without formal authority. - Stack: ours is PySpark on EMR, Hudi/Delta, Airflow, with SQL and Python everywhere. Deep in something comparable, fast ramp on the rest. - Operations: you can deploy, debug, and un-break distributed workloads in the cloud without waiting for another team. - Nice to have: streaming (Kafka or Kinesis); healthcare, claims, or other regulated-data experience. - Experience using AI-assisted coding tools (e.g., GitHub Copilot, Claude Code) to accelerate development while maintaining quality is encouraged COMPENSATION This role pays $190,000 to $230,000 per year, based on experience, in addition to stock options. Anticipated role close date: 9/15/2026