Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Staff Site Reliability Engineer

Lightspeedhq
CompanyLightspeedhq
CategoryEngineering
LocationMontreal
RemoteHybrid
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted24 Jul 2026
Last verified30 Jul 2026
SourceEmployer career page (ashby)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
Hi there! Thanks for stopping by 👋 Are you actively looking for a new opportunity? Or just checking the market? Well… you might just be in the right place! We’re looking for a Staff Site Reliability Engineer to join our Data team in Canada. As a Staff Data SRE, you are the technical backbone of the Data Office's infrastructure platform. Your scope spans the entire Data Business Unit — you solve systemic, platform-level problems, not individual tickets. You own the reliability, scalability, and developer experience of the data platform, and you act as a force multiplier: making the engineers around you faster, the systems more resilient, and the platform easier to consume. You bring deep GCP expertise and a product mindset to data infrastructure. You are comfortable navigating ML/AI workload requirements (Vertex AI, feature stores, training pipelines) and can reason confidently about the BI tooling and data layers that feed it (Looker, BigQuery). Your decisions carry weight across teams, and you communicate them clearly to both engineers and non-technical stakeholders. Please note this role is based in Montreal, Canada What you’ll be doing: - You will own and be accountable for advancing the Data Office infrastructure engineering practices and delivering high-leverage platform projects. Your time will be distributed as follows: - 70% – Hands-On Engineering & Platform Ownership: Design and deliver significant infrastructure improvements. You write production-quality IaC (Terraform), lead solutions from design through delivery, and are the person the team calls for the hardest platform problems. This includes data infrastructure for batch/streaming workloads, ML/AI environments (Vertex AI, model serving, GPU-backed compute), and the BI serving layer (Looker infrastructure and GCP integration). - 10% – Team Coordination & Technical Leadership: Lead architecture and design discussions for larger cross-team projects. You review critical PRs, drive solution design sessions, and provide technical direction that keeps the team coherent and moving forward. - 10% – General Meetings: Lead conversations in agile ceremonies, incident reviews, and cross-functional syncs. You shift from participant to driver. - 10% – Mentorship & Technical Review: Actively mentor Senior and Intermediate SREs. You set the bar for IaC quality, observability practices, and operational discipline through reviews, pairing, and documentation. - Avoiding Pitfalls: You proactively identify risks in complex data migrations or infrastructure changes and propose mitigation strategies to ensure zero data loss and minimal downtime. And a little bit of… - Participating in on-call rotation and incident response. - Contributing as part of the wider team to achieve organization-wide objectives even if this means doing things that aren't strictly within the scope of your role. What you need to bring: - GCP Infrastructure: Deep expertise across GCP compute, networking, IAM, GKE, data services, and FinOps. - IaC Proficiency: Terraform as primary tool. - Hands-on experience with Looker infrastructure and ML/AI platform tooling (Vertex AI, model serving, training pipelines). - Proficient in Bash and Golang; Python a plus for data tooling. - Strong experience with metrics, logs, traces, alerting, and SLO/SLI design, DataDog,... - Product Mindset: You focus on making the platform easy to use, not just "available." - Experience with Github actions, Circle CI, GCP Cloud Build, … - Ability to communicate infrastructure trade-offs to both engineers and non-technical stakeholders. - AI proficiency, Go-to AI expertise within the Data BU — evaluates tooling, drives cross-team AI-first development practices. - Security-First Mindset, every design decision evaluated through a security and compliance lens. - Self-awareness with a willingness to learn and improve. - Ability to mentor and tr
HOUSE AD970,107 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →