Volta is a French–Italian AI-powered platform that digitises and automates B2B sales operations for suppliers and distributors.
ABOUT THE ROLE Volta builds and operates large scale GPU compute infrastructure for AI workloads. The network is not a layer underneath the platform, it is part of it. Fabric design, overlay and multi-tenancy, edge connectivity, and the software that programs and observes all of it sit in one team, deliberately. You will lead that team: a small group of senior network platform engineers who write production software against the fabric rather than only operating it. WHAT YOU WILL BE DOING - Lead the network platform engineering team: technical direction, design reviews, code review, and day-to-day delivery. - Set the technical direction for the network layers of the platform: compute and storage fabric, overlay and multi-tenant isolation, edge connectivity, and the telemetry and automation that make them operable. - Own the design and evolution of Volta's fabric standards across sites, including RoCE v2 and InfiniBand deployments, and hold the team to a single reference architecture rather than per site variants. - Represent network platform engineering in roadmap planning: translate product requirements into scoped, sequenced technical work, communicate trade-offs, and keep delivery on track. - Own how network capability is exposed upward: the APIs, abstractions, and IaaS control plane integration through which tenants get isolated, performant networking. - Work closely with the bring-up teams to surface operational pain points from cluster deployment and turn them into scalable platform features and automation. - Evaluate and challenge network designs from OEMs and partners, and hold vendors to Volta's requirements rather than accepting reference designs as given. - Support customer facing technical conversations: workload requirements, fabric design review, and acceptance criteria. - Collaborate with the security engineering team on trust boundaries, tenant isolation, and remediation of network layer findings. - Coordinate with the other platform engineering leads on shared architecture, joint initiatives, and cross-team and cross-timezone delivery. - Own reliability, observability, and interface and versioning standards for the services and fabric your team ships. - Stay hands-on: take on complex or high-risk engineering work directly alongside the team, including during bring-up and incident response. - Run post-mortems and close structural gaps after incidents, not just the immediate issue. - Manage the team: 1:1s, performance conversations, hiring interviews, and onboarding. WHAT YOU BRING - 5+ years in large scale data center or cloud network engineering, with at least 2 years leading an engineering team. - Production software development, not scripting. Python or Go in a shared repository under normal review and CI standards. Our working languages are Python, Go, and Rust. - Ethernet fabric design and operations at depth: leaf spine, BGP including unnumbered BGP, ECMP, and the day-two realities of running it at low latency. - RoCE v2 at scale: PFC and ECN tuning, DCQCN, and a clear understanding of how it behaves differently from InfiniBand in a GPU training environment. - InfiniBand production experience: fat tree topology, UFM, fabric partitioning, adaptive routing, congestion control, with operational and troubleshooting depth rather than design familiarity alone. - EVPN and VXLAN in production, including running an overlay under multi-tenant load. - Network automation at scale: configuration as code, a source of truth system such as NetBox or Nautobot, CI validation of network change, and declarative or idempotent tooling. - Kubernetes networking: CNI, and how workload networking interacts with the underlying fabric. - GPU infrastructure experience: how training and inference traffic patterns shape fabric design, and how capacity is surfaced through a platform or API layer. - Multi-vendor credibility: able to design, evaluate, and challenge proposals from any major OEM. - Security awareness: trust boundaries, least privilege, and secure defaults in multi-tenant network design. - Proven people management: running 1:1s, delivering performance feedback, and taking accountability for a team's delivery and wellbeing. - Clear communicator who can translate technical complexity for product, leadership, and customer stakeholders. - Willing to be on site during cluster bring-up when it matters. NICE TO HAVE (BUT NOT ESSENTIAL) None of these are required. Several map to specific areas of the team's scope, so strength in one or more helps. - Fluency with AI-assisted development, and interest in scaling agent-assisted workflows across the team (agentic CLI tools, MCP, skills, APIs) to amplify delivery. - ASN operations: running a public autonomous system, BGP peering with transit providers and IXes, RPKI and IRR hygiene, and DDoS posture. - IPv6 at production scale: dual stack DC design, v6 BGP peering, and addressing architecture. - NVIDIA Spectrum-X, including NetQ and Cumulus, or SONiC and whitebox platforms. - gNMI, OpenConfig, or NETCONF/YANG based telemetry and configuration. - OVN/OVS, SR-IOV, DPDK, or BlueField DPU based networking. - Familiarity with NVLink and NVSwitch topologies and NCCL behaviour. - Experience working distributed across time zones with counterparts in other regions. - Open source contributions to networking or infrastructure projects.