Senior Research Engineer
at Turing · 51-100 employees
- Seniority
- Senior
- Location
- San Francisco, California, United States
- Posted
- 5d ago
at Turing · 51-100 employees
Turing is a startup building fully autonomous driving systems powered by an end-to-end AI stack and a large-scale foundation model.
<div class="content-intro"><h1><span style="font-family: helvetica, arial, sans-serif;"><strong><span style="font-size: 24pt;">About Turing</span></strong></span></h1> <p><span id="m_-1033525666082252048m_-8288224233410703305gmail-docs-internal-guid-39386146-7fff-394b-7c89-9e58d80485b6">Turing’s mission is to accelerate superintelligence to drive real economic progress. Headquartered in San Francisco, Turing works with frontier AI labs to generate high-quality datasets, reinforcement learning environments, and frontier research benchmarks that improve model capabilities in software engineering, enterprise knowledge work, and advanced STEM reasoning. In software engineering, Turing is the largest and longest-running data provider in the category. Turing also works with Fortune 500 enterprises across financial services, life sciences, healthcare, retail, automotive, and CPG to build and deploy end-to-end agentic AI systems inside mission-critical workflows. By operating on both sides, Turing closes the loop between frontier research and enterprise deployment, turning real-world deployment signals into better data, evaluations, and more capable models. Learn more at <a href="http://www.turing.com/" target="_blank" data-saferedirecturl="https://www.google.com/url?q=http://www.turing.com&source=gmail&ust=1788032823652000&usg=AOvVaw2LLfe9HeuAcQz5ewoGP-0C">www.turing.com</a>. </span></p> <div> </div></div><h2>The Role</h2> <p>Turing builds large-scale datasets and reinforcement learning (RL) environments that power post-training for the world’s leading AI labs and enterprises, including OpenAI, Anthropic, Google DeepMind, Microsoft AI, Amazon, Apple, and many more. We create RL environments to evaluate and improve our customers' models on complex, long-range, multi-step workflows across high-GDP-value domains such as Finance, Sales, Retail, Developer Tools, Collaboration, Customer Experience. </p> <p>The environments vary depending on the model capability being evaluated / improved, a few examples of environment types are listed here: </p> <ol> <li>Environments for Software Engineering / coding agents </li> <li>UI-Environments for Computer-Use/Browser-Use agents </li> <li>MCP-based Environments for general function-calling agents across various enterprise and consumer applications</li> </ol> <p>The <strong>Senior Research Engineer</strong> will own end-to-end the creation of datasets, RL environments, and evals for frontier AI labs in the domain of coding agents and software engineering. This is a <strong>hands-on technical leadership role</strong> where you influence revenue directly – you will be mapped to one or more AI labs and interface directly with researchers / engineers at those labs to understand their needs and build data offerings to address those needs. To achieve this, you will build and manage teams of software engineers, researchers, QAs, and contractors/data-annotators from Turing’s talent pool of 4M+ developers. </p> <p>You’ll be responsible for delivering projects at frontier quality and scale—owning data quality, throughput, and timely delivery. You’ll define and manage data pipelines, validation workflows, and review processes to ensure datasets meet the highest standards for realism, correctness, and diversity. You’ll also develop automations, synthetic data generation systems, and internal tools to scale production efficiently. </p> <p>In short, you’ll run your project like a startup within Turing, owning both the technical architecture and the operational execution required to produce best-in-class datasets/environments/evals to make the world’s best coding agents and models even better at real-world coding tasks across the software development lifecycle.</p> <h2>What you’ll do</h2> <p><strong style="font-size: 14px;">1. End-to-End Ownership: Data Quality, Process Design, and Team Building</strong></p> <ul> <li>Lead the creation of datasets, rl environments, and evals focused on Coding Agents / Software Engineering for one or more AI lab customers.</li> <li>Ensure that everything you ship to clients meets frontier standards for realism, correctness, diversity, and difficulty.</li> <li>Set up quality rubrics, automated validation scripts, and human review processes for every stage of data generation.</li> <li>Build and lead cross-functional teams of software engineers, researchers, QAs, and data creators drawn from Turing’s 4M+ developer network.</li> <li>Interview, onboard, train, and mentor team members to ensure consistent output quality and technical excellence.</li> </ul> <p><strong>2. Collaborate with Researchers at Frontier Labs</strong></p> <ul> <li>Act as the primary technical point of contact for your customer projects, interfacing directly with researchers and engineers at frontier AI labs to understand their coding agent roadmap and model data needs, to gather feedback, and to co-define success criteria for your projects.</li> <li>Provide regular progress updates, surface insights from model evaluations, and incorporate client feedback to improve future iterations.</li> </ul> <p><strong>3. Drive Research, Sales Enablement, and Industry Thought Leadership</strong></p> <ul> <li>Fine-tune models in-house on Turing-generated datasets or Turing-rl-environment generated trajectories to determine model improvement as a proof of data quality</li> <li>Proactively build benchmarks and run evals on frontier models and coding agents to identify strengths and weaknesses on SWE tasks, and leverage these insights to inform product roadmap</li> <li>Equip customer-facing teams with the Evaluation reports, sample datasets, and trainings to enable them to communicate your data offerings to customers most effectively</li> <li>Publish research papers and technical posts on Turing’s data products, innovations in our synthetic data generation / automation pipelines, evaluations of frontier agents and models, an