Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Senior Network Engineer - InfiniBand / UFM

Lightning AI
CompanyLightning AI
CategoryEngineering
LocationNew York City
RemoteOn-site (inferred)
EmploymentNot stated
LevelSenior
SalaryUSD 170k–210k
Posted20 Jul 2026
Last verified12 Aug 2026
SourceThe employer's own careers page (company_site)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
Lightning AI, the company behind PyTorch Lightning, is seeking a Senior Network Engineer to design, deploy, and operate high-performance NVIDIA InfiniBand fabrics that power large-scale GPU clusters for AI training and inference. This role is critical to scaling Voltage Park's AI Factory infrastructure from tens of thousands of GPUs to significantly beyond, ensuring deterministic performance and enterprise-grade reliability. What You'll Do • Design, deploy, and maintain large-scale NVIDIA InfiniBand fabrics supporting AI/ML GPU clusters • Deploy and administer NVIDIA Unified Fabric Manager (UFM) Enterprise for monitoring, provisioning, telemetry, and fabric health • Configure and optimize NVIDIA Quantum and Quantum-2 InfiniBand switches, including firmware lifecycle management • Troubleshoot fabric performance issues impacting NCCL, MPI, GPUDirect RDMA, and AI training jobs using packet captures and diagnostic tools • Automate network provisioning using Python, Ansible, Git, REST APIs, and Infrastructure-as-Code methodologies; monitor health using UFM telemetry, Prometheus, and Grafana What You Need • 7+ years of data center networking experience • 3+ years supporting NVIDIA InfiniBand environments with hands-on experience deploying and operating Quantum and Quantum-2 InfiniBand switches • Strong Linux administration experience (Ubuntu) and deep understanding of Layer 2 and Layer 3 networking • Experience with automation using Python and Ansible; expertise with BGP, EVPN, VXLAN, and spine-leaf architectures • Proficiency with packet captures and troubleshooting using tcpdump, Wireshark, and ibdiagnet tools; strong understanding of InfiniBand Architecture including Subnet Manager, Adaptive Routing, Congestion Control, Partition Keys, LIDs, Queue Pairs, Virtual Lanes, and Service Levels Nice to Have • 10+ years of large-scale data center networking experience • Experience with Netris and Terraform • Experience designing networks for hyperscalers, neoclouds, or high-scale SaaS infrastructure • Exposure to bare-metal provisioning systems and multi-region backbone design • Experience working in high-growth infrastructure startups Anticipated annual base salary range: $170,000–$210,000 USD. Total rewards package includes discretionary bonus, meaningful equity component, comprehensive medical/dental/vision coverage, retirement and financial wellness support, generous paid time off, paid parental leave, professional development support, wellness and work-from-home stipends, and flexible work environment.