Senior DevOps / Infrastructure Engineer
Category Labs
| Company | Category Labs |
| Category | Engineering |
| Location | New York |
| Remote | Hybrid |
| Employment | Not stated |
| Level | Senior |
| Salary | USD 180k–250k |
| Posted | 5 Aug 2026 |
| Last verified | 12 Aug 2026 |
| Source | Employer ATS (ashby) |
Description
Category Labs (formerly known as Monad Labs) is a team of systems engineers and researchers on a mission to design and build at the frontier of decentralized technology. We strive to deliver significant improvements over existing blockchain solutions. After raising $225M in series A funding, led by Paradigm, we are growing our team.
We’re the team behind Monad, a high-performance, EVM-compatible Layer 1 whose public mainnet is now live. We write the core software that runs it: a parallel-execution EVM https://github.com/category-labs/monad, a custom state database, and a BFT consensus client https://github.com/category-labs/monad-bft, all developed in the open.
THE ROLE
We're looking for a Senior DevOps / Infrastructure Engineer to operate the infrastructure behind Monad, and to push how much of that operation can be driven by AI. You'll keep our globally-distributed validator, full node, and archive fleet healthy across mainnet and testnet, own our infrastructure-as-code and observability, and build the agentic tooling and guardrails that let a small team safely operate a large fleet. As more of our engineering shifts toward AI, this role is central to designing the workflows, deterministic guardrails, and security perimeters within which autonomous agents operate our infrastructure; you'll also help stand up and operate the infrastructure behind our own growing model workloads.
WHAT YOU'LL DO
- Operate the Monad node fleet: health, sync, upgrades, and recovery across validators, full nodes, archive/historical, and indexer nodes on mainnet and testnet, including safe, staged rollouts and incident response.
- Own our infrastructure-as-code: Ansible for fleet configuration, Terraform + Atlantis for cloud and DNS, and Kubernetes/Flux (GitOps) for platform services.
- Build and operate observability and alerting (Prometheus, Grafana, Loki); create dashboards and alerts that catch problems before they page while minimizing false positives.
- Automate the release pipeline: node upgrades, canary rollouts, snapshot/restore, and the guardrails that bound blast radius (e.g., protecting validators from automated changes).
- Design and build agentic operations: develop AI agents, tooling (e.g., MCP servers), and runbooks-as-code that let agents safely investigate, diagnose, and execute routine operations, with deterministic guardrails and human oversight.
- Codify operational knowledge into tools and automation that the whole team, and its agents, can reuse.
- Harden nodes and services, manage secrets, and continuously drive down manual toil.
WHO YOU ARE
- You have 5+ years in DevOps, SRE, or Infrastructure Engineering, operating production systems at scale.
- You have strong Linux, systemd, networking, and shell fundamentals, and you're comfortable debugging live systems over SSH.
- You have deep, hands-on infrastructure-as-code experience with Ansible and Terraform.
- You have experience with observability stacks (Prometheus, Grafana, Loki, or equivalents).
- You have hands-on fluency with AI-assisted engineering: you use coding agents and LLM tooling in your daily workflow and have judgment on where it helps and where it's risky.
- You have experience designing automation with safe guardrails, and you bring calm, methodical incident response.
- You have programming and scripting experience (e.g., Python, bash).
- Experience with Kubernetes and GitOps (Flux or Argo) is a plus.
- Experience building AI agent tooling, MCP servers, or agent orchestration frameworks is a plus.
- Experience serving inference, either locally or as a service is a plus.
- Previous experience with blockchain clients or node operations is a plus.
- A Bachelor of Science in Computer Science, Engineering, or a related field is a plus.
WHY WORK WITH US
- Challenging problems: You’ll work on extremely challenging problems with massive impact. See our Blogs https://www.category.xyz/blogs and Publications &