Senior DevOps / Infrastructure Engineer
Category Labs
| Company | Category Labs |
| Category | Engineering |
| Location | New York City |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Senior |
| Salary | USD 180k–250k |
| Posted | 6 Aug 2026 |
| Last verified | 12 Aug 2026 |
| Source | The employer's own careers page (company_site) |
Description
Category Labs is building Monad, a high-performance EVM-compatible Layer 1 blockchain with a live public mainnet. This Senior DevOps / Infrastructure Engineer role focuses on operating the globally-distributed validator and node fleet, managing infrastructure-as-code, and designing AI-assisted operations tooling that allows a small team to safely manage large-scale distributed systems.
What You'll Do
• Operate the Monad node fleet across mainnet and testnet, ensuring health, synchronization, upgrades, and recovery for validators, full nodes, archive nodes, and indexers with staged rollouts and incident response
• Own infrastructure-as-code using Ansible for fleet configuration, Terraform + Atlantis for cloud and DNS management, and Kubernetes/Flux for GitOps-driven platform services
• Build and operate observability and alerting infrastructure using Prometheus, Grafana, and Loki; create dashboards and alerts that detect issues proactively while minimizing false positives
• Automate the release pipeline including node upgrades, canary rollouts, and snapshot/restore operations with guardrails to protect validators and limit blast radius
• Design and build agentic operations tooling: develop AI agents, MCP servers, and runbooks-as-code that enable autonomous investigation, diagnosis, and routine operations execution with deterministic guardrails and human oversight
• Codify operational knowledge into reusable automation tools accessible to the engineering team and AI agents; harden nodes and services, manage secrets, and continuously reduce manual operational toil
What You Need
• 5+ years of hands-on experience in DevOps, SRE, or Infrastructure Engineering operating production systems at scale
• Strong Linux, systemd, networking, and shell scripting fundamentals with demonstrated ability to debug live systems over SSH
• Deep, hands-on infrastructure-as-code experience with both Ansible and Terraform
• Experience with observability stacks such as Prometheus, Grafana, Loki, or equivalent tools
• Hands-on fluency with AI-assisted engineering including regular use of coding agents and LLM tooling in daily workflow with judgment about appropriate applications and risks
• Experience designing automation with safe guardrails, calm and methodical incident response, and programming/scripting experience in Python, Bash, or similar languages
Nice to Have
• Experience with Kubernetes and GitOps tools (Flux or Argo)
• Experience building AI agent tooling, MCP servers, or agent orchestration frameworks
• Experience serving inference either locally or as a service
• Previous experience with blockchain clients or node operations
• Bachelor of Science in Computer Science, Engineering, or a related field
Base salary $180,000–$250,000 across US locations plus equity and token incentives. Full-time benefits include private health insurance, flexible paid time off, monthly wellness reimbursement, paid parental leave. US employees receive 100% paid medical, dental, vision insurance (75% dependent coverage), HSA/FSA options, 401(k) with company match, and lunch/dinner stipend for NYC office. Non-US employees receive EOR-based benefits and country-specific packages.