Data Engineer - Data Platform
at Mill
- Location
- San Bruno, California
- Posted
- 5d ago
at Mill
<div class="content-intro"><p><span style="font-weight: 400;">Mill is a waste prevention technology company reimagining what it means to eliminate waste, starting with food. We build smart systems and infrastructure for homes, businesses, and municipalities that transform food scraps from landfill-bound waste into valuable resources, including chicken feed. Tens of thousands of Mill’s residential food recyclers are already helping households divert millions of pounds of food scraps every year, paving the way for our upcoming launch of Mill Commercial—the industry’s first end-to-end solution for managing, understanding, and preventing food waste in commercial environments (e.g. grocery, restaurants, food services). At Mill, we are passionate about building easy-to-use, beautifully designed technologies that keep food in the food system and out of landfills.</span></p></div><h2><span style="font-weight: 400;">The Role</span></h2> <p><span style="font-weight: 400;">As a Data Engineer at Mill, you'll build and maintain the core data infrastructure that powers analytics and product data across the company — ingestion pipelines, warehouse modeling, data quality, and the self-serve analytics platform (Hex + Snowflake) our business teams rely on. You'll work closely with the Senior Data Engineer owning our recommendations platform, contributing to and supporting that work as needed, with the opportunity to grow into deeper recommendation/LLM-based work over time. You'll partner closely with product, engineering, data analytics, and marketing teams.</span></p> <h2><span style="font-weight: 400;">What You'll Do</span></h2> <ul> <li style="font-weight: 400;"><span style="font-weight: 400;">Design, build, and maintain scalable data pipelines across Mill's product and operational systems</span></li> <li style="font-weight: 400;"><span style="font-weight: 400;">Manage and maintain data infrastructure that powers our product and operational systems, ensuring it's reliable and ready to feed external customers and downstream analytics</span></li> <li style="font-weight: 400;"><span style="font-weight: 400;">Collaborate with software engineers to instrument new product features and ensure event data flows cleanly </span></li> <li style="font-weight: 400;"><span style="font-weight: 400;">Help build and maintain the self-serve analytics platform in Hex and Snowflake for internal business teams</span></li> <li style="font-weight: 400;"><span style="font-weight: 400;">Help maintain the metrics, table endorsements, and business logic that analysts and stakeholders rely on</span></li> <li style="font-weight: 400;"><span style="font-weight: 400;">Own data quality monitoring — build alerting, validation frameworks, and observability tooling so data issues get caught before they become business problems</span></li> </ul> <h2><span style="font-weight: 400;">What We're Looking For</span></h2> <ul> <li style="font-weight: 400;"><span style="font-weight: 400;">3-5 years of experience operating data engineering systems in production</span></li> <li style="font-weight: 400;"><span style="font-weight: 400;">Have built and operated data pipelines in production using Python, SQL, and tools like dbt, Airflow, Fivetran, or similar against a cloud data warehouse (e.g., Snowflake)</span></li> <li style="font-weight: 400;"><span style="font-weight: 400;">Have used Infrastructure as Code (e.g., Terraform, Pulumi) to provision and manage data infrastructure, with CI/CD discipline for pipeline and infra changes (automated testing, staged rollout, rollback)</span></li> <li style="font-weight: 400;"><span style="font-weight: 400;">Experience working with transactional databases (e.g., PostgreSQL, Amazon RDS) as a data source, including understanding how OLTP systems differ from warehouse/analytical workloads</span></li> <li style="font-weight: 400;"><span style="font-weight: 400;">Comfort working in a collaborative environment where data consumers are partners, not just stakeholders, and comfort moving between different types of work as priorities shift</span></li> <li style="font-weight: 400;"><span style="font-weight: 400;">A bias toward action</span></li> </ul> <h2><span style="font-weight: 400;">Nice to Have</span></h2> <ul> <li style="font-weight: 400;"><span style="font-weight: 400;">Exposure to recommendation, personalization, or LLM-based product logic — not required, but a strong plus given the team's direction</span></li> <li style="font-weight: 400;"><span style="font-weight: 400;">Experience building or supporting self-serve analytics tooling (Hex, Looker, or similar)</span></li> <li style="font-weight: 400;"><span style="font-weight: 400;">Exposure to distributed systems concepts (partitioning, consistency, fault tolerance)</span></li> <li style="font-weight: 400;"><span style="font-weight: 400;">Experience with Mixpanel, Tableau, or similar BI/analytics tools</span></li> <li style="font-weight: 400;"><span style="font-weight: 400;">Familiarity with data contract or data mesh patterns, or RBAC/access governance on a warehouse</span></li> <li style="font-weight: 400;"><span style="font-weight: 400;">Experience with event tracking or product analytics</span></li> </ul> <p><span style="font-weight: 400;"><em data-stringify-type="italic">The estimated base salary range for this position is $185k to $210k, <em>which does not include the value of benefits or a potential equity grant. A wide range of factors are considered in making compensation decisions, including but not limited to skill sets, market conditions, experience and training, licensure and certifications, and business and organizational needs.</em></em></span></p>