Type: Full-time permanent contract or an Internship with a potential follow-up offer Location: San Francisco or remote with future relocation to San Francisco (sponsored) Start: ASAP ABOUT US - We're building state-of-the-art context compression. Our mission is to become the "Cloudflare for LLMs" — a compression layer embedded into most LLM pipelines by default. - We're a team of ex-EPFL MSc/PhDs from dlab. We started by publishing papers, then got into YC and started making money helping companies cut their LLM costs. - We run the business like a research lab: form hypotheses, kill the ones that don't work, double down on the ones that do. ABOUT YOU: - A cracked full-stack engineer who enjoys a high-paced startup environment, takes pride in what they build and owns it end to end. - Python/basic ML Ops skills. Experience in scaling AI infra products is a plus. - Proactive, strong communicator with fast response time, team player TECH REQUIREMENTS - Strong backend engineering fundamentals - Experience with concurrency and distributed systems - Ability to work across systems (Python + light frontend) - Excellent Claude Code (or similar) user NICE TO HAVE - Open-source contributions - Startup experience - OAuth / API auth flows STACK - Backend: Python, FastAPI, PostgreSQL (Supabase), Redis, AWS - Frontend: Next.js, React, TypeScript - Tools: GitHub, Docker, Sentry, GitHub Actions INTERVIEW PROCESS 1. Intro call (20 min) 2. Practical technical interview (60 min) 3. Cultural interview (30 min)