Staff Software Engineer, Data Products
at Omada Health
- Seniority
- Staff Principal
- Work model
- Remote
- Location
- Remote, USA
- Posted
- 15d ago
at Omada Health
<p>Omada Health is on a mission to bend the curve of chronic disease. We rely on trusted, high-quality data to power intelligent products, personalized member experiences, and data-driven decision making.</p> <p>As machine learning becomes increasingly central to our platform, we're investing in the data foundations that enable scalable model development, experimentation, and production inference.</p> <p><strong>Job overview:</strong><strong><br></strong></p> <p>We are seeking a <strong>Staff Software Engineer, Data Engineering</strong> to lead the design and development of the production data platform that powers machine learning across Omada.</p> <p>In this role, you will partner closely with <strong>Data Scientists, Applied AI Engineers, Product Engineers, and fellow Data Engineers</strong> to identify, design and build trusted, reusable datasets foundations that serve as the foundation for feature engineering, model training, experimentation, and production inference.</p> <p>Rather than building one-off pipelines for individual models, you'll create scalable data products and feature pipelines that enable multiple machine learning use cases while ensuring consistency, reliability, and governance across the ML lifecycle.</p> <p>You will own the technical design of feature datasets—from ingesting raw behavioral, clinical, and operational data through transforming, validating, and publishing production-grade datasets that are reusable across modeling teams.</p> <p>This role is ideal for someone who enjoys solving complex data problems, designing scalable distributed data systems, and enabling machine learning through well-engineered data foundations.</p> <p><strong>Key Responsibilities:</strong></p> <ul> <li>Design, build, and maintain reusable feature datasets that support machine learning use cases including personalization, engagement, risk prediction, churn modeling, recommendation systems, and experimentation. </li> <li>Establish self-service foundations that streamline and democratize dataset creation across the data organization.</li> <li>Partner with Data Scientists to translate modeling requirements into production-ready feature pipelines, supporting the full model lifecycle from exploration to deployment.</li> <li>Identify source data, transformations, and historical windows needed for feature engineering. Help define and build shared, reusable feature definitions across models rather than one-off datasets.</li> <li>Balance features freshness, correctness, latency, and computational efficiency when designing data pipelines.</li> <li>Build datasets that support both historical model training and future production inference.</li> <li>Design and implement batch and streaming pipelines that transform raw healthcare, behavioral, product, and operational data into trusted ML-ready datasets.</li> <li>Build reliable data processing systems using Python, SQL, Spark, and modern cloud data platforms.</li> <li>Optimize large-scale distributed processing for performance, scalability, and cost.</li> <li>Design data pipelines that are modular, testable, observable, and easy to evolve as product requirements change.</li> <li>Ensure data quality through testing, anomaly detection, schema validation, and pipeline monitoring.</li> <li>Partner with platform teams to support near real-time feature generation where appropriate.</li> <li>Improve reproducibility by standardizing feature computation across experimentation and production.</li> <li>Support rapid experimentation without sacrificing long-term maintainability.</li> <li>Ensure data quality through testing, anomaly detection, schema validation, and pipeline monitoring.</li> <li>Establish engineering standards for correctness, documentation, and maintainability.</li> <li>Familiarity with feature stores or feature management platforms.</li> <li>Familiarity with model training pipelines and MLOps workflows.</li> </ul> <p><strong>Technical Leadership:</strong></p> <ul> <li>Lead architecture and design discussions for large-scale ML data systems, driving adoption of reusable patterns and platform capabilities across Data Engineering.</li> <li>Influence technical direction across multiple engineering teams, embedding with Product, Engineering, and business stakeholders (Clinical, Finance, Growth, Enrollment) during early design phases to shape data capture requirements at the source.</li> <li>Translate ambiguous business requirements from Business domain SMEs into concrete technical specs, maintaining consistency of business logic and definitions across systems.</li> <li>Mentor engineers on distributed data processing, software engineering best practices, and scalable data modeling.</li> </ul> <p><strong>About you: </strong></p> <p><strong>Experience</strong></p> <ul> <li>8+ years building large-scale production data platforms and distributed data pipelines.</li> <li>Experience designing reusable datasets that power machine learning, experimentation, or advanced analytics.</li> <li>Demonstrated experience partnering closely with Data Scientists to productionize feature engineering workflows.</li> <li>Experience leading cross-team technical initiatives and influencing engineering direction.</li> <li>Strong experience working with cloud-native data platforms such as AWS.</li> <li>Experience building production data systems using Databricks, Iceberg, Spark, Redshift, Snowflake, or similar technologies.</li> <li>Experience developing reliable batch and streaming data pipelines.</li> <li>Experience working with healthcare, behavioral, or other large-scale event data is a plus.</li> </ul> <p><strong>Technical Skills</strong></p> <ul> <li>Expert SQL with strong data modeling skills.</li> <li>Strong programming skills in Python, Java, or Scala.</li> <li>Experience with Apache Spark or similar distributed compute frameworks.</li> <li>Experience with Airflow or similar orchestration platforms.</li> <li>Experience designing dimensional models, even