Traversal builds AI agents that autonomously detect, troubleshoot and resolve production incidents for enterprise-grade site reliability engineering.
ABOUT TRAVERSAL Traversal https://www.traversal.com/ is the AI Site Reliability Engineer (SRE) for the enterprise—already trusted by some of the largest companies in the world to troubleshoot, remediate, and even prevent the most complex production incidents. Our mission is to free engineers from endless firefighting and enable them to focus on creative, high-impact work. Our roots remain deeply embedded in AI research, and we’re channeling that scientific rigor and creativity into building the premier AI agent lab for the enterprise. Hence, what we’re proudest of is assembling the most talented yet nicest group of individuals, including researchers from MIT, Harvard, and Berkeley, to world-class engineers from industry: Citadel Securities, Cockroach Labs, Datadog, DE Shaw, ServiceNow, Glean, Perplexity, Pinecone, and more, to take on one of the hardest problems for AI to solve. Without the entire team, none of this would be possible. THE ROLE Our core product is our AI SRE agent. As a Product Engineer on the Agents team, you'll own how that agent behaves, where it shows up, and whether it holds up under real production load. That means working end-to-end on the harness and context engineering that shape agent reasoning, the streaming and orchestration layers that put it in front of engineers mid-incident, and the observability and evaluation systems that tell us whether any of it is actually working. Three areas drive impact on this team: - Agent Behavior: Improving how the agent reasons, decides, and communicates through harness design and context engineering. This is prompt and tool architecture, not prompt tweaking: what the agent sees, when it acts, what it does with ambiguity, and how we measure whether a change made it better. - Live Surface Area: Getting agent intelligence into the places engineers already work, in real time. Slack channels, the web app, and whatever comes next. Streaming, orchestration, and long-running work that survives restarts and reconnects. - Stability at Scale: Our agents run for hours against noisy, high-volume customer telemetry. Making that reliable means observability into agent behavior, scalable data patterns, and systems that degrade gracefully instead of silently. Each of these requires product vision and autonomy. You'll be deciding what to build as often as you're building it. RESPONSIBILITIES - Agent Development: Build AI agents with a mind for user experience and capabilities like long-horizon work, proactiveness, resilience, readability, and evidence citation. - User Experience: Innovate and drive the evolution of AI UX/UI design, rapidly iterating based on user feedback to improve overall usability. - Harness and Context Engineering: Design the tool interfaces, context assembly, and decision logic that determine how well the agent performs on real incidents. - Evaluation: Build the eval harnesses and offline datasets that let us ship agent changes with confidence rather than vibes. - Live Delivery: Design and implement the streaming, orchestration, and API layers that carry agent output into Slack and the product in real time. - Data and Infrastructure: Own efficient storage, retrieval, and processing across PostgreSQL, Redis, Kafka, and S3, in support of agents operating on large-scale time-series and topological data. - Product Impact: Translate agent capability into features that reduce cognitive load for on-call engineers, iterating quickly on user feedback. REQUIREMENTS - 3+ years of software engineering experience, with a strong focus on full-stack or backend systems. - Hands-on experience building with agentic AI systems or LLM-powered products, including familiarity with tool use, context management, and the failure modes that come with both. - Strong experience with Python and web frameworks such as FastAPI. - Comfort working in a React and TypeScript codebase, enough to ship a feature end to end without waiting on someone else for the frontend. - Experience deploying applications on AWS, working with ECS or Kubernetes, Postgres for data storage, and S3 for large-scale object storage. - Proven ability to lead in fast-paced startup environments, with limited resources, shifting priorities, and minimal structure. You don't need to meet every requirement. If you have deep experience in a subset of these areas, we encourage you to apply. NICE TO HAVE - Experience with durable execution or workflow orchestration frameworks such as Temporal. - Background in observability, monitoring, or incident response tooling, as a builder or a heavy user. - Experience running evals for LLM systems, or otherwise bringing measurement discipline to non-deterministic software. - Background in large-scale, complex, data-driven applications. COMPENSATION We offer competitive compensation, startup equity, health insurance, and additional benefits. The U.S. base salary range for this full-time, in-person role in New York is $150,000–$300,000, plus equity and benefits. Our salary ranges are based on location, level, and role. Individual compensation is determined by experience, skills, and job-related knowledge. WHY YOU SHOULD JOIN US We’ll make sure you’re fully supported with health insurance, a great tech setup, flexible time off, and plenty of in-office snacks. We offer competitive salary and equity packages, and take thoughtful consideration with every hire on our small, high-impact team. Traversal is fully in-office, 5 days a week, based in New York near Madison Square Park. We have a collaborative, hard-working culture and are energized by building the future of AI-powered software maintenance. Working here means owning meaningful parts of the product, having the flexibility to move fast, and learning constantly. This is a place to grow your career, make a real impact, and help define a new category of infrastructure software.