Faster chat, better deals — Get the App

SRE Manager (Site Reliability Engineering)

Indeed

Company

Job typeFull-time
Workplace typeOnsite
Experience levelNo experience limit
Education levelNo degree limit

Description

Job Summary: The SRE Manager leads the team of reliability engineers, ensuring system and infrastructure availability, performance, and efficiency. Key Highlights: 1. Lead the SRE team, defining the team's vision and roadmap 2. Ensure reliability and performance of large-scale environments 3. Drive automation with IaC, CI/CD, and observability Job Description: The SRE Manager (Site Reliability Engineering) is responsible for leading the site reliability engineering team, ensuring system and infrastructure availability, performance, and efficiency. Requirements: * Prior experience in SRE or systems operations roles * Knowledge of system monitoring and analysis tools * Leadership and teamwork skills * Ability to efficiently solve complex problems * Familiarity with DevOps practices and automation * Excellent communication and documentation skills * Academic background in Information Technology or related fields **Experience** * Solid experience in Software Engineering, DevOps, or SRE, with proven leadership/management experience over technical teams. * History of working in high-scale and mission-critical environments. **Technical Stack** * Public cloud: AWS, GCP, Azure, and/or OCI. * Orchestration and containers: Kubernetes and Docker. * IaC: Terraform and Ansible. * Observability: Prometheus, Grafana, Datadog, and OpenTelemetry. * Programming languages: Go, Python, and Bash. * Relational and NoSQL databases in production. **Practices** * Disaster Recovery, infrastructure security, and FinOps. * Incident management and blameless postmortem culture. * Capacity planning and performance optimization in microservices environments. **Soft Skills** * Leadership and people development. * Clear communication with both technical and executive audiences. * Decision-making under pressure and analytical reasoning. * Business acumen and prioritization capability. Responsibilities **Strategic Leadership** * Define the SRE team’s vision, roadmap, and processes, aligning infrastructure with business objectives. * Establish and evolve SLI, SLO, SLA, and error budget policies. * Represent the SRE function to stakeholders across Product, Engineering, Security, and Business teams. **People and Project Management** * Lead, develop, and engage the SRE team, conducting 1:1s, PDIs, performance reviews, and hiring processes. * Manage Capex and Opex budgets, tracking KPIs for efficiency and cost. * Prioritize initiatives and balance short-term deliveries with long-term reliability investments. **Reliability and Performance** * Improve environment resilience by monitoring availability, latency, error rate, and saturation. * Plan infrastructure capacity, performance, and costs (FinOps). * Lead Disaster Recovery initiatives, resilience testing, and chaos engineering. **Incident Response** * Actively manage crises and critical incidents (serve as incident commander when required). * Promote a blameless post\-mortem culture and continuous learning. * Establish healthy and sustainable on\-call routines for the team. **Automation and Platform Engineering** * Drive adoption of IaC (Terraform, Ansible) and CI/CD pipelines. * Reduce toil through automation and building reusable platform capabilities. * Establish observability standards (logs, metrics, traces) using Prometheus, Grafana, Datadog, and OpenTelemetry. Work Model: On-site \- Recife \- PE

Some content was automatically translated

Posted by

João Silva

Indeed · HR

Location

João Silva

Indeed · HR

Similar jobs

SRE Manager (Site Reliability Engineering) by Indeed in 2026 | ok.com