Infrastructure Operations Engineer (APAC)
Lightning AI
| Company | Lightning AI |
| Category | Engineering |
| Location | SG |
| Remote | Remote |
| Employment | Not stated |
| Level | Not stated |
| Salary | SGD 165k–205k |
| Posted | 27 Jul 2026 |
| Last verified | 12 Aug 2026 |
| Source | The employer's own careers page (company_site) |
Description
Lightning AI, the company behind PyTorch Lightning, is seeking an experienced Infrastructure Operations Engineer to join the APAC InfraOps team. This role is central to reliability, automation, and operational scale for GPU infrastructure, handling break/fix operations, incident response, customer provisioning, observability, and automation systems for large-scale infrastructure.
What You'll Do
• Design, build, and roll out new platforms and patterns to minimize incidents and enable customer-facing and internal features
• Deploy updates and improvements to support Voltage Park's internal and end-customer use cases
• Collaborate with Infrastructure Engineering, Network Operations, Customer Success, and Software Platform Development teams on troubleshooting and operational efficiency
• Participate in on-call rotation (evenly distributed across team members in primary/secondary pattern) to respond to incidents and manage GPU infrastructure reliability
• Build automation systems that reduce manual toil and improve operational efficiency across large-scale GPU environments and bare metal infrastructure
What You Need
• 8+ years working with Linux as a server/hosting platform
• 5+ years experience with AWS
• 2+ years experience with Kubernetes and strong container fundamentals
• 2+ years experience with Terraform and Ansible
• 2+ years with network attached storage management (NFS, Ceph, or equivalent protocols)
• Experience with monitoring systems (Prometheus, ELK stack)
• Familiarity with GitOps workflow
• Software development experience using Python, Go, Bash, or other languages for automation and system integration
• Deep networking fundamentals including datacenter-level networking knowledge
• Experience building and delivering complex systems with ability to navigate tradeoffs between design, risk, cost, and outcomes
• Strong written and oral communication skills
Nice to Have
• Experience with bare metal hardware troubleshooting and provisioning, particularly Dell hardware
• Experience with GPU servers in bare metal or virtualized form
• Deep experience with network switches, routers, and firewalls (SONiC, Palo Alto, Juniper Networks)
• Experience with VAST storage systems
$165,000–$205,000 SGD annually. Total rewards include discretionary bonus, meaningful equity (RSUs), comprehensive health coverage (medical, dental, vision), 401(k) matching (U.S.) and pension contributions (U.K.), unlimited PTO, company holidays, paid parental and family leave, annual learning and development allowance, wellness and work-from-home stipends, four weeks paid sabbatical after four years of service, flexible schedules, and complimentary meals at office hubs.