Specialist Cloud Site Reliability Engineer
NICE
| Company | NICE |
| Category | Engineering |
| Location | Pune |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 29 Jul 2026 |
| Last verified | 12 Aug 2026 |
| Source | The employer's own careers page (company_site) |
Description
NICE is seeking a Senior Site Reliability Engineer to join the core Reliability Engineering team, responsible for ensuring the scalability, reliability, and performance of mission-critical systems and observability platforms across multiple environments and regions. This role is ideal for someone who thrives in fast-paced environments, enjoys automation, and has a strong background in cloud-native operations, observability stacks, and incident management.
What You'll Do
• Design and implement scalable, reliable, and resilient systems across hybrid or multi-cloud environments (primarily AWS/EKS/ECS/Lambda)
• Build and manage infrastructure automation using Terraform, Helm, and Kubernetes; improve CI/CD pipelines using Jenkins and GitHub Actions
• Own and enhance the observability stack including Prometheus, Grafana, Loki, Tempo, and OpenTelemetry; define and implement SLOs and error budgets
• Lead major incident response, root cause analysis, and blameless postmortems; partner with product teams on operational readiness
• Mentor junior SREs and developers on reliability practices, automation, and observability
What You Need
• 8+ years of strong experience with Kubernetes, EKS, ECS, and containerized workloads in production
• Expertise in AWS services including EC2, Lambda, IAM, RDS, S3, ALB/NLB, VPC, and PrivateLink
• Proficiency with Terraform, Helm, Jenkins, GitHub Actions, and GitOps tools such as ArgoCD or Flux
• Deep understanding of observability frameworks covering metrics, logs, traces, and distributed monitoring
• Strong knowledge of Linux, networking fundamentals, and system performance tuning
Nice to Have
• Hands-on experience with Prometheus, Grafana, Loki, Tempo, Alloy, OpenTelemetry, or equivalent observability tools
• Familiarity with Python, Go, or Shell scripting for automation and custom tooling
• Practical experience in incident response, root cause analysis, and on-call operations