Data Engineer, Analytics Data Products
- Work model
- Hybrid
- Location
- New York, NY
- Posted
- 6d ago
<div class="content-intro"><div id="labeledImage.LOCATION" class="WIFG" data-automation-id="decorationWrapper"> <div class="WNHJ"> <div id="labeledImage.LOCATION--uid38" class="WE-Y WMXY WBAB WF0Y" data-automation-id="responsiveMonikerInput" data-metadata-id="labeledImage.LOCATION" data-uxi-form-item-child-list-index="0"> <div class="WJ-Y"> <p><strong>The <a href="https://www.nytco.com/mission-and-values/" target="_blank"><u>mission</u></a> of The New York Times is to seek the truth and help people understand the world. That means independent journalism is at the heart of all we do as a company. It’s why we have a world-renowned newsroom that sends journalists to report on the ground from nearly 160 countries. It’s why we focus deeply on how our readers will experience our journalism, from print to audio to a world-class digital and app destination. And it’s why our business strategy centers on making journalism so good that it’s worth paying for. </strong></p> </div> </div> </div> </div></div><h3 data-pm-slice="1 3 []"><strong>About the Role:</strong></h3> <p>At The New York Times, data powers decisions across the entire company. The Analytics Data Products team builds the foundational data products and pipelines that make that possible, and we're looking for a <strong>Data Engineer</strong> to help us build them. You'll own and enhance the data pipelines and core, reusable data products that partner teams across the company depend on to unlock analytics for their most important questions. You'll work hands-on across our hybrid cloud architecture (AWS and GCP) and contribute to the platform that delivers trusted data products company-wide spanning multiple business domains. You'll join a collaborative team that invests in your growth. Reporting to the Senior Engineering Manager of Analytics Data Products, you'll take ownership of pipelines and products, learn from experienced engineers, and grow your impact across the company.</p> <p>This is a hybrid role in our New York City headquarters.</p> <h3><strong>Responsibilities:</strong></h3> <ul> <li> <p>Design, model, and implement complex data pipelines for the cleansed and curated data layers in the medallion architecture, taking full ownership of the data product's structure, partitioning, documentation, and performance characteristics.</p> </li> <li> <p>Develop advanced data transformations using dbt (data build tool) for relational data modeling and PySpark for complex data processing within the Lakehouse, ensuring outputs meet strict SLAs and quality standards.</p> </li> <li> <p>Collaborate with Data Analysts and other consumers to define requirements and translate them into scalable data models suitable for analytic use cases.</p> </li> </ul> <ul> <li> <p>Manage physical data storage across both GCP (GCS, BigQuery, Cloud Composer) and AWS (S3, Glue, Athena, EMR).</p> </li> <li> <p>Choose optimal file formats such as Parquet and Iceberg, and design efficient partitioning and clustering strategies.</p> </li> <li> <p>Administer and tune Spark compute<em> </em>resources (e.g., Dataproc, EMR, or managed services) to optimize job execution time and cost.</p> </li> <li> <p>Optimize user queries and access patterns to maintain platform performance and cost efficiency.</p> </li> <li> <p>Implement centralized data quality checks and observability mechanisms within the data pipeline to proactively identify and resolve data issues.</p> </li> <li> <p>Contribute to the implementation of metadata management, data lineage, and role-based access control (RBAC) programs across the Lakehouse environment.</p> </li> </ul> <h3>Basic Qualifications</h3> <ul> <li> <p><strong>2+ years</strong> of full-time professional, hands-on experience with Software Engineering in a data context or equivalent experience</p> </li> <li> <p>Strong proficiency in <strong>Python</strong> for scripting and data manipulation</p> </li> <li> <p><strong>Strong proficiency in SQL</strong> and demonstrable experience with complex, <strong>production-level data modeling</strong> (preferably dimensional modeling, Kimball, OBT, or Data Vault)</p> </li> <li> <p>Demonstrated experience <strong>owning data pipelines and products end-to-end </strong>through the full SDLC</p> </li> <li> <p>Hands-on experience with a <strong>Cloud Data Warehouse</strong> (<strong>BigQuery, Snowflake, DataBricks</strong>)</p> </li> <li> <p>Familiarity with foundational cloud services and data storage components in at least one major cloud provider (<strong>GCP or AWS</strong>)</p> </li> <li> <p>Experience with workflow orchestration tools (e.g., Airflow, Cloud Composer, or Prefect) and version control systems (<strong>Git</strong>)</p> </li> </ul> <h3>Preferred Qualifications</h3> <ul> <li> <p>Experience operating in a dual-cloud environment (GCP/AWS)</p> </li> <li> <p>Experience with Infrastructure-as-Code (IaC) tools like Terraform</p> </li> <li> <p>Knowledge of <strong>PySpark</strong> or other Spark APIs</p> </li> <li> <p>Experience with advanced Lakehouse file formats like <strong>Iceberg or Delta Lake</strong></p> </li> <li> <p>Familiarity ensuring data product SLAs and quality standards, integrating advanced testing, quality checks, and monitoring into the CI/CD pipeline</p> </li> </ul> <p><span style="font-weight: 400;">REQ-019488</span></p> <p><span style="font-weight: 400;">#LI-hybrid</span></p> <p><span style="font-weight: 400;"><span data-sheets-root="1"><strong>Additional Compensation and Benefits:<br><br></strong></span></span><span style="font-weight: 400;"><span data-sheets-root="1">For roles in the U.S., dependent on your role, you may be eligible for variable pay, such as an annual bonus and restricted stock. Benefits may include medical, dental and vision benefits, Flexible Spending Accounts (F.S.A.s), a company-matching 401(k) plan, employee stock purchase plan, paid vacation, paid sick days, paid parental leave, tuition reimbursement and professional development programs. <br><b