Tebra provides an integrated EHR+ platform that streamlines clinical, billing, telehealth, and marketing workflows for private healthcare practices.
<div class="content-intro"><p>Tebra only initiates contact with candidates via email from an official Tebra email address (@<a class="c-link" href="http://tebra.com/" target="_blank" data-stringify-link="http://tebra.com" data-sk="tooltip_parent">tebra.com</a>, @<a class="c-link" href="http://patientpop.com/" target="_blank" data-stringify-link="http://patientpop.com" data-sk="tooltip_parent">patientpop.com</a>, or @<a class="c-link" href="http://kareo.com/" target="_blank" data-stringify-link="http://kareo.com" data-sk="tooltip_parent">kareo.com</a>) or through our applicant tracking system, Greenhouse. We will only ask you to provide sensitive personal information through our official application portal — not via social media or text message. We do not conduct interviews via instant messaging.</p></div><h1>About the Role</h1> <p>As a <strong>Data Engineer focused on AI/ML</strong>, you’ll build, maintain, and optimize the data infrastructure that powers Tebra’s intelligent features. You’ll partner closely with Machine Learning Engineers, Data Scientists, and Software Engineers to transform complex healthcare data into high-quality datasets and real-time features that enable machine learning models.</p> <p>This is a hands-on engineering role where you’ll contribute to scalable data pipelines, improve data quality, and help ensure our AI systems are powered by reliable, performant, and well-governed data. You’ll work on modern data platforms and gain experience building solutions that support both model training and production inference.</p> <h1>Your Area of Focus</h1> <ul> <li><strong>Design, build, and maintain</strong> scalable data pipelines for feature extraction, training data generation, and model monitoring.</li> <li>Develop and enhance data systems that support analytics and machine learning workloads, including data lakehouse and feature store technologies.</li> <li>Monitor production data pipelines, identify data quality issues or pipeline failures, and implement improvements to ensure reliability and freshness.</li> <li>Participate in engineering design discussions and contribute to technical decisions around data architecture and pipeline implementation.</li> <li><strong>Build reusable data engineering components</strong>, including automated data quality checks, schema validation, and testing frameworks.</li> <li>Translate business requirements into scalable data solutions that enable analytics and machine learning use cases.</li> <li><strong>Optimize SQL queries, Spark workloads, and data processing pipelines </strong>to improve performance and scalability.</li> <li>Collaborate with <strong>ML Engineers</strong> and cross-functional partners to support MLOps best practices, including data versioning, lineage, and reproducibility.</li> <li>Break down technical work into manageable tasks and deliver high-quality solutions within an agile team.</li> </ul> <h1>Your Professional Qualifications</h1> <ul> <li><strong>3+ years </strong>of professional experience in Data Engineering, Software Engineering, or a related field.</li> <li>2+ years of hands-on experience building and maintaining production data pipelines supporting analytics, reporting, or machine learning workloads.</li> <li>Strong proficiency in <strong>Python</strong> and SQL with experience developing production-quality data pipelines.</li> <li>Experience with modern data processing technologies such as <strong>Spark, Airflow, Kafka,</strong> or similar distributed data platforms.</li> <li>Experience working with cloud-based data platforms such as <strong>Databricks</strong>, <strong>Snowflake</strong>, <strong>Delta Lake</strong>, or equivalent <strong>lakehouse</strong> technologies.</li> <li>Understanding of data modeling, data warehousing, and data governance best practices.</li> <li>Familiarity with machine learning data workflows, including training datasets, feature engineering, and data quality concepts.</li> <li>Experience deploying and supporting production data pipelines with monitoring, testing, and <strong>CI/CD</strong> practices.</li> <li>Strong problem-solving skills, attention to detail, and the ability to collaborate effectively across engineering and product teams.</li> <li>Excellent communication skills and a desire to continuously learn new technologies and engineering practices.</li> </ul> <p><strong>#LI-SS1 #LI-Remote</strong></p><div class="content-pay-transparency"><div class="pay-input"><div class="description"><p>We are dedicated to attracting and retaining top talent with competitive and fair compensation. For this position, this range reflects our Zone 1 (National Average) pay band. Your specific compensation is thoughtfully determined by your experience, qualifications, the specific requirements of the role, and your Geo Zone. Our geo-zone system ensures your pay is competitive for your location, recognizing varying costs of labor across regions.</p> <p>Our four geo zones are designed to reflect this:<br>Zone 1: National Average<br>Zone 2: Moderately Higher Cost Regions<br>Zone 3: High-Cost Regions<br>Zone 4: Lower-Cost Regions</p> <p>Beyond base compensation, Tebra offers eligible employees the opportunity for variable pay and a robust benefits package, reflecting our commitment to your overall well-being. In compliance with California pay transparency laws, the specific compensation range applicable to your Geo Zone will be shared during your initial talent screen.</p></div><div class="title">Zone 1 (National Average)</div><div class="pay-range"><span>$128,000</span><span class="divider">—</span><span>$145,200 USD</span></div></div></div><div class="content-conclusion"><p> </p> <h1>About Tebra</h1> <p>Tebra is the only all-in-one EHR+ platform built exclusively for independent healthcare practices. Designed to replace the clunky, fragmented tools built for corporate systems, Tebra connects EHR software, billing, automation, telehealth solution, and marketing — so providers ca