Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Staff Software Engineer, Cloud Monitoring Service (Distributed Systems)

Crusoe
CompanyCrusoe
CategoryEngineering
LocationSan Francisco
RemoteOn-site (inferred)
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted8 Jul 2026
Last verified9 Aug 2026
SourceEmployer ATS (ashby)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster. We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI. We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services. If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe. About the Role: We are seeking a Staff Software Engineer to lead distributed systems design and development within Crusoe Cloud's Cloud Monitoring Service. This team builds the observability backbone of Crusoe Cloud: the telemetry systems that collect, process, store, and serve the metrics and logs our customers depend on to run AI workloads at scale. You will own the design and evolution of high-throughput, multi-tenant data systems, from collection at the edge through ingestion, storage, and query. This is a full-time position. What You'll Be Working On: - Distributed Systems Ownership: Own the architecture and evolution of large-scale telemetry pipelines, including high-volume ingestion, stream processing, time-series and log storage, and low-latency query paths. Design systems that stay correct, available, and cost-efficient as data volume grows 10x. - Scalable, Multi-Tenant Design: Design services that are highly scalable, durable, and fair across tenants. Solve the hard problems in this space: hot shards, high-cardinality data, noisy neighbors, backpressure, retention and compaction at scale, and graceful degradation under load. - Reliability and Operational Excellence: Build for operability from day one. Improve pipeline reliability and data freshness, reduce on-call burden through better system design rather than more process, and participate in a customer-facing on-call rotation, leading by example. - Technical Leadership: Set the technical direction for the team's distributed systems work. Drive design reviews, identify one-way door decisions early, and raise the bar on how the team scopes, builds, and operates systems. - Cross-Team Collaboration: Work with product, compute, networking, and platform teams to make sure observability decisions are made with full context. Represent the team's technical position in cross-org conversations. - Mentorship: Coach senior and mid-level engineers through design work, code review, and incident response. Build patterns and frameworks that make the team better without requiring your direct involvement. What You'll Bring to the Team: - Distributed Systems Depth: Deep, hands-on experience designing and operating distributed systems at scale. You have solved real problems in sharding, replication, consistency, load balancing, and concurrency, not just studied them. - Observability Data Experience: Experience building or operating large-scale data infrastructure such as time-series databases, log aggregation, streaming pipelines, or distributed tracing backends. Familiarity with technologies like Prometheus, VictoriaMetrics, Loki, OpenTelemetry, Kafka, Vector, or similar. - Technical Proficiency: Strong programming fundamentals in Go or another modern compiled language (Go strongly prefer