TrueFoundry provides a Kubernetes-based, cloud-agnostic PaaS that automates ML workflows—model deployment, autoscaling, monitoring, RAG/agent workflows, GPU/memory optimization, and centralized access controls for multi-cloud and on-premises environments.
<p><strong>About TrueFoundry</strong><br><br>Every production AI system, whether it's powering customer support, writing code, analyzing financial data, or diagnosing medical conditions, needs the same foundational infrastructure.A way to route between models. A way to manage tools and integrate them securely. A way to orchestrate agents and enforce governance. A unified compute layer to run it all.</p> <p><strong>That infrastructure layer is being built right now.</strong></p> <p>We are looking for a Staff/Principal Engineer – Core Team.</p> <h2><strong>The Problem We're Solving</strong></h2> <p>Companies are moving beyond simple chatbots to production agentic systems. These systems route between OpenAI, Anthropic, Google, and self-hosted models. They integrate dozens of tools via protocols like MCP. They orchestrate multi-agent workflows where agents coordinate with other agents.</p> <p>The infrastructure to support this doesn't exist yet. You can't just duct-tape together a few API calls and call it production-ready.</p> <p>You need a control plane that handles:</p> <ul> <li>Intelligent routing with observability, cost policies, and fallback logic</li> <li>Centralized tool and MCP server management with security and lifecycle controls</li> <li>Agent orchestration with governance and guardrails</li> <li>A unified compute layer to run self-hosted models, custom tools, and agents</li> </ul> <p><strong>AI Gateway</strong> is the control plane: five composable components (Prompts, LLM Gateway, MCP Gateway, Guardrails, Agent Gateway) that handle routing, orchestration, and governance.</p> <p>We're Series A, backed by Intel Capital and Sequoia. Companies like CVS, Mastercard, Siemens, Paytm, Synopsys, and Zscaler run production AI workloads on our platform.</p> <h3><span style="font-size: 14px;"> </span><strong>The Role:</strong></h3> <ul> <li>Solve some of the most complex Engineering problems and drive it alongside a team of engineers & ML researchers.</li> <li>Build a deep, holistic understanding of the TrueFoundry platform across all components and shape the product vision and implementation.</li> <li>Act as the technical face of engineering for customer-related discussions and escalations</li> <li>Guide and unblock engineers across projects in the US region</li> <li>Partner closely with our CTO and India-based engineering team to drive system design, architecture, and implementation of complex products</li> <li>Lead technical design, critical customer problem-solving, and platform scalability initiatives end-to-end</li> </ul> <p>This is a high-ownership, high-impact role designed for an engineer who loves combining world-class systems thinking with real-world execution.</p> <h3><strong>What You’ll Do:</strong></h3> <ul> <li data-section-id="pu1709" data-start="142" data-end="273">Build and scale TrueFoundry's MCP Gateway and Agentic Gateway, enabling secure, reliable, and scalable AI agent interactions.</li> <li data-section-id="15vrwnd" data-start="274" data-end="398">Design and develop cloud-native, distributed systems that power AI agents, tool integrations, and enterprise AI workloads.</li> <li data-section-id="2qg1e3" data-start="399" data-end="555">Build core capabilities such as authentication, authorization, routing, observability, security, and multi-tenant infrastructure for the gateway platform.</li> <li data-section-id="13ytixl" data-start="556" data-end="675">Collaborate closely with the CTO to shape the technical architecture, product roadmap, and long-term platform vision.</li> <li data-section-id="1jz8w31" data-start="676" data-end="807">Drive architectural decisions, participate in design and code reviews, and ensure high engineering standards across the platform.</li> <li data-section-id="1vneuh5" data-start="808" data-end="952">Work closely with enterprise customers to understand their AI infrastructure needs and translate feedback into scalable platform capabilities.</li> <li data-section-id="1ygruid" data-start="953" data-end="1054">Mentor engineers across teams, helping them build high-quality, reliable, and maintainable systems.</li> <li data-section-id="pc4xtp" data-start="1055" data-end="1181">Continuously improve platform performance, reliability, scalability, and developer experience while reducing technical debt.</li> <li data-section-id="3pz3r" data-start="1182" data-end="1324" data-is-last-node="">Stay at the forefront of emerging AI infrastructure, Model Context Protocol (MCP), and agentic systems, bringing new ideas into the product</li> </ul> <h3><strong>Who You Are:</strong></h3> <ul> <li>8+ years of strong backend/systems engineering experience at top technology companies or startups</li> <li>Deep expertise in distributed systems, cloud-native architectures, and scalable system design</li> <li>Strong working knowledge of Kubernetes, containerized workloads, and infrastructure engineering</li> <li>Practical experience building or deploying ML/GenAI applications (or closely working with ML/DS teams)</li> <li>Skilled in programming languages such as Python, Golang, Node.js, and TypeScript</li> <li>Solid understanding of system observability, resiliency design, and SRE practices</li> <li>Strong technical leadership and communication skills — able to work with both customers and engineering teams</li> <li>Ability to think strategically while also executing hands-on when required</li> </ul> <p><strong>Traits we are looking for:</strong> Ownership, ability to execute, hustle and think out of the box, data-driven decision making, be comfortable with more unknowns than knowns.</p> <h3>Perks of Working at TrueFoundry</h3> <ul> <li>Join a fast-growing Series A, Bay Area-based startup building cutting-edge AI infrastructure.</li> <li>Comprehensive health insurance for you and your family, including medical, dental, and vision coverage.</li> <li>401(k) retirement plan.</li> <li>Flexible hybrid work, 2 days a week in the office (Tuesday & Wednesday