Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Lead DevOps Engineer

Telus Digital
CompanyTelus Digital
CategoryEngineering
LocationRemote
RemoteRemote
EmploymentNot stated
LevelLead
SalaryNot stated by the employer
Posted22 Jun 2026
Last verified30 Jul 2026
SourceEmployer career page (ashby)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
LEAD DEVOPS ENGINEER, SITE RELIABILITY WHO WE ARE Welcome to TELUS Digital https://www.telusdigital.com/— where innovation drives impact at a global scale. As an award-winning digital product consultancy and the digital division of TELUS https://www.telus.com/en/, one of Canada’s largest telecommunications providers, we design and deliver transformative customer experiences through cutting-edge technology, agile thinking, and a people-first culture. With a global team across North America, South America, Central America, Europe, Africa, and APAC, we offer end-to-end expertise across eight core service areas: Digital Product Consulting, Digital Marketing Services, Data & AI, Strategy Consulting, Business Operations Modernization, Enterprise Applications, Cloud Engineering, and QA & Test Engineering. From mobile apps and websites to voice UI, chatbots, AI, customer service, and in-store solutions, TELUS Digital enables seamless, trusted, and digitally powered experiences that meet customers wherever they are — all backed by the secure infrastructure and scale of our multi-billion-dollar parent company. Location and Flexibility This role can be fully remote for candidates based in Brazil, due to team distribution and occasional in-person opportunities. If you are based in São Paulo or Porto Alegre, you are welcome to work from one of our offices on a flexible schedule. About the Role Our CXAI Platform powers a portfolio of Generative AI products deployed into enterprise contact centers and BPO operations, environments where downtime, latency, or silent model degradation translate directly to commercial impact. As a Staff DevOps Engineer, Site Reliability, you'll lead the architecture and maintenance of the infrastructure and reliability practices that keep AI-powered systems performant, observable, and trustworthy under real production load, including redundancy, latency, and cost management. This is a staff-level individual contributor role with broad mandate. You'll set technical standards across the platform, partner directly with product and engineering leadership, and have real ownership over how reliability shapes the roadmap. WHAT YOU'LL OWN - Platform reliability strategy: help define SLOs/SLIs for AI-powered services, including latency and quality SLOs for LLM inference paths, and build the error-budget discipline that lets product teams ship fast without breaking trust. - Cloud architecture on GCP: design scalable, secure infrastructure for distributed AI services, event-driven workloads, and multi-LLM-provider integrations - Observability for non-deterministic systems: build metrics, tracing, and alerting that surface not just "is it up" but "is it behaving correctly" for LLM-powered features (drift, regression, hallucination rates, tool-call failures) - Resilience engineering: circuit breakers, graceful degradation, multi-provider failover, and chaos/fault-injection practices for AI inference paths - Infrastructure-as-code and automation: Terraform-first, automated everything, no toil tolerated - Production readiness: define and enforce PRR-style standards across teams launching new AI products and features - Technical leadership: mentor engineers, drive architecture reviews, and shape the broader engineering culture around reliability WHAT YOU BRING - Significant infrastructure engineering experience combining DevOps and SRE disciplines at scale - Deep GCP expertise (AWS a strong plus); relevant cloud certifications welcome - Production experience with SRE fundamentals: SLO/SLI design, error budgets, toil reduction, blameless incident review - Strong background in distributed systems failure modes and resilience patterns - Expert-level infrastructure-as-code (Terraform), container orchestration (Kubernetes), and CI/CD - Hands-on with modern observability stacks (i.e., OpenTelemetry, Sentry) and AI-specific observability tooling (Arize, LangSmith, Braintrust, or simila
HOUSE AD976,467 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →