Zócalo Health is a tech-enabled, community-based primary care provider delivering integrated medical, behavioral, and social care to underserved, high-need populations.
<p><strong>Senior Data Engineer</strong></p> <p>at Zócalo Health </p> <p>Remote (Full Time) </p> <p>Compensation: $160,000 - $180,000 (per year)</p> <p> </p> <p><strong>About Us</strong></p> <p>Zócalo Health is a tech-enabled, community-oriented primary care organization serving people who have historically been underserved by the one-size-fits-all healthcare system. We partner with health plans, providers, and community organizations to deliver culturally competent primary care, behavioral health, and social care.</p> <p>Our model is built for populations with high medical and social complexity, where fragmented care drives poor outcomes and unnecessary cost. We combine local, community-based teams with virtual care and modern technology to deliver coordinated, whole-person care where members live and receive support.</p> <p>Founded in 2021, Zócalo Health is backed by leading healthcare and mission-aligned investors and is scaling rapidly across states and populations. We are building a durable care platform designed to perform in constrained healthcare environments and to lead the shift toward accountable, value-based care.</p> <p> </p> <p><strong>Role Description</strong></p> <p>The <strong>Senior Data Engineer</strong> will join Zócalo Health as we build the data platform that powers analytics, product measurement, and operational visibility across the company. This is a hands-on building role at a foundational stage: you will design and ship the pipelines, ingestion frameworks, and data models that the rest of the company depends on.</p> <p>The primary focus of this role is establishing a scalable, durable data platform. This includes laying the groundwork for longer-term initiatives such as the longitudinal patient record, population-level analytics, and product instrumentation. You will partner closely with Engineering and Product to ensure the data platform supports roadmap priorities and outcome measurement as the company grows.</p> <p>This position reports to the <strong>Principal Data Engineer</strong> and partners closely with Engineering and Product.</p> <p> </p> <p><strong>In your first 12 months, you will:</strong></p> <ul> <li>Build and operate production-grade ingestion pipelines from core clinical, operational, and third-party systems into our Databricks lakehouse</li> <li>Develop and maintain dbt models that turn raw data into clean, well-documented, analytics-ready datasets</li> <li>Establish data quality, testing, and monitoring practices that make pipelines reliable and trustworthy</li> <li>Help shape ingestion patterns and architecture standards alongside the Principal Data Engineer</li> <li>Enable company-wide metrics for care outcomes and operations</li> <li>Collaborate with cross-functional leads to develop and iterate on a suite of core operational dashboards, ensuring teams have the self-service tools they need to track company metrics and outcomes.</li> </ul> <p> </p> <p><strong>The Senior Data Engineer will contribute in the following ways:</strong></p> <ul> <li>Design, build, and operate production data pipelines across clinical, operational, and third-party systems using API-based ingestion, Change Data Capture (CDC), and event- or webhook-driven patterns</li> <li>Build and maintain transformation layers in dbt, including tests, documentation, and reusable models</li> <li>Develop and refine core analytical and longitudinal data models used across the company</li> <li>Implement testing, monitoring, and observability to ensure data quality, pipeline reliability, and system performance</li> <li>Apply strong engineering fundamentals to improve the scalability, performance, and cost-efficiency of data systems on AWS and Databricks</li> <li>Partner with Product to support metric definitions, outcome measurement, and reporting needs</li> <li>Contribute to engineering standards, code review, and a culture of knowledge sharing and continuous improvement</li> <li>Partner with business, product, and engineering stakeholders to design and build intuitive data visualizations and dashboards that drive actionable insights and program visibility.</li> </ul> <p> </p> <p><strong>Core Technologies (current and planned)</strong></p> <ul> <li>Cloud: AWS</li> <li>Lakehouse / data platform: Databricks</li> <li>Transformations: dbt</li> <li>Languages: SQL and Python (primary languages for ingestion and transformation)</li> <li>Ingestion patterns: API-based ingestion, Change Data Capture (CDC), and event- or webhook-driven pipelines, including frameworks such as PySpark and Spark Structured Streaming on Databricks</li> <li>Orchestration: workflow orchestration (e.g., Databricks Workflows or Airflow)</li> </ul> <p> </p> <p><strong>Qualifications</strong></p> <ul> <li>5+ years of experience in data or backend engineering roles with significant data platform responsibility</li> <li>Hands-on experience building and operating production-grade data pipelines and ingestion frameworks</li> <li>Strong proficiency in SQL and Python for data ingestion, processing, and transformation</li> <li>Experience with a cloud data platform; experience with AWS and Databricks (or a comparable Spark-based lakehouse) strongly preferred</li> <li>Experience building SQL-based transformation workflows; hands-on experience with dbt preferred</li> <li>Strong computer science fundamentals, including comfort reasoning about distributed systems and data processing at scale</li> <li>Ability to diagnose and resolve performance, reliability, and data quality issues in complex systems</li> <li>Strong ownership mindset and comfort operating in ambiguous, fast-growing environments</li> <li>Clear communicator able to partner effectively with technical and non-technical stakeholders</li> <li>Experience building dashboards or analytical outputs used by executives and frontline teams</li> </ul> <p> </p> <p><strong>Preferred Qualifications</strong></p> <ul> <li>Experience working