Upscale AI is a Santa Clara–based startup building synchronized AI-cluster networking that integrates GPUs, accelerators, storage and networking.
Why join Upscale AI Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads. We focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact. If you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard. Location: India · Experience: 12–15 years · Type: Full-time, IC About us We build a distributed control plane for datacenter network fabrics. Our platform orchestrates intent-driven configuration, change management, drift remediation, and full-stack observability across cloud-hosted services and on-prem edge appliances deployed in customer datacenters worldwide. The role You will own critical platform subsystems end-to-end — from design through production. You'll build and ship core distributed systems components, drive the technical quality of the codebase, and unblock the team through direct hands-on contribution. You work closely with Product, QA, and Customer Engineering to deliver quality product to stakeholders. What you'll work on Distributed control plane spanning cloud services and on-prem edge appliances connected via mTLS gRPC streams Observability and telemetry at scale: OpenTelemetry collection pipelines, stream processing (Kafka), time-series storage (ClickHouse/VictoriaMetrics), real-time fabric state views, packet-event analysis, and fleet-wide health aggregation Agentic AI operations: design and build autonomous infrastructure agents that evaluate prerequisites, orchestrate multi-step workflows (onboarding, upgrades, drift remediation), handle failure recovery, and interact with the control plane through tool-use patterns (LangGraph, MCP) Fleet orchestration: aggregate health/compliance/drift APIs, cross-site campaign execution, template promotion workflows, parallel site onboarding Data architecture across Postgres, ArangoDB, ClickHouse, Redis, Kafka, and Git-backed content stores Multi-tenant SaaS with site-scoped RBAC, session lifecycle, and enterprise IdP integration Device lifecycle management: zero-touch provisioning, enrollment protocols, config push via edge relay, drift detection and remediation Intent compilation engine: workspace management, merge request lifecycle, change request execution with gate-based verification and rollback What you bring 12–15 years building and operating distributed systems serving enterprise customers across cloud and on-prem environments Deep proficiency in Go (or comparable systems language) with strong distributed systems fundamentals Significant experience designing observability and telemetry platforms: collection agents, stream processing, time-series databases, alerting pipelines, and real-time dashboards at scale Production experience with microservices architecture: gRPC, Protocol Buffers, spec-first REST APIs (OpenAPI) Hands-on with multiple storage paradigms: relational, graph, time-series, key-value, and streaming Track record building multi-tenant platforms with tenant isolation, RBAC, and identity federation Experience with Kubernetes, Helm, and hybrid cloud/on-prem deployment models Strong API design sense: versioning, backward compatibility, contract-first development Ability to communicate architectural decisions clearly through writing and diagrams How we expect you to work Cross-functional by default — you work closely with Product, QA, Design, and Customer Engineering, not in isolation Solution-oriented — when the team is stuck, you unblock them by driving toward answers and building what's needed Accountable for delivery — you take personal responsibility for shipping quality product to stakeholders on time Hands-on always — you write code daily and prove ideas by building them Nice to have AI/ML agent architectures for infrastructure operations: LangGraph, AutoGen, MCP tool-use, human-in-the-loop gating, and autonomous workflow orchestration Network automation or infrastructure management platforms (Apstra, NSO, Terraform, Crossplane) Datacenter networking experience: OpenConfig, gNMI, or fabric management at scale Hub-and-spoke / edge computing / control-plane-data-plane separation architectures OpenTelemetry contributor experience or deep familiarity with the collector ecosystem Device enrollment or zero-touch provisioning systems Open-source contributions or published work in distributed systems