Description
Job Summary:
The SRE Manager leads the team of reliability engineers, ensuring system and infrastructure availability, performance, and efficiency.
Key Highlights:
1. Lead the SRE team, defining the team's vision and roadmap
2. Ensure reliability and performance of large-scale environments
3. Drive automation with IaC, CI/CD, and observability
Job Description:
The SRE Manager (Site Reliability Engineering) is responsible for leading the site reliability engineering team, ensuring system and infrastructure availability, performance, and efficiency.
Requirements:
* Prior experience in SRE or systems operations roles
* Knowledge of system monitoring and analysis tools
* Leadership and teamwork skills
* Ability to efficiently solve complex problems
* Familiarity with DevOps practices and automation
* Excellent communication and documentation skills
* Academic background in Information Technology or related fields
**Experience**
* Solid experience in Software Engineering, DevOps, or SRE, with proven leadership/management experience over technical teams.
* History of working in high-scale and mission-critical environments.
**Technical Stack**
* Public cloud: AWS, GCP, Azure, and/or OCI.
* Orchestration and containers: Kubernetes and Docker.
* IaC: Terraform and Ansible.
* Observability: Prometheus, Grafana, Datadog, and OpenTelemetry.
* Programming languages: Go, Python, and Bash.
* Relational and NoSQL databases in production.
**Practices**
* Disaster Recovery, infrastructure security, and FinOps.
* Incident management and blameless postmortem culture.
* Capacity planning and performance optimization in microservices environments.
**Soft Skills**
* Leadership and people development.
* Clear communication with both technical and executive audiences.
* Decision-making under pressure and analytical reasoning.
* Business acumen and prioritization capability.
Responsibilities
**Strategic Leadership**
* Define the SRE team’s vision, roadmap, and processes, aligning infrastructure with business objectives.
* Establish and evolve SLI, SLO, SLA, and error budget policies.
* Represent the SRE function to stakeholders across Product, Engineering, Security, and Business teams.
**People and Project Management**
* Lead, develop, and engage the SRE team, conducting 1:1s, PDIs, performance reviews, and hiring processes.
* Manage Capex and Opex budgets, tracking KPIs for efficiency and cost.
* Prioritize initiatives and balance short-term deliveries with long-term reliability investments.
**Reliability and Performance**
* Improve environment resilience by monitoring availability, latency, error rate, and saturation.
* Plan infrastructure capacity, performance, and costs (FinOps).
* Lead Disaster Recovery initiatives, resilience testing, and chaos engineering.
**Incident Response**
* Actively manage crises and critical incidents (serve as incident commander when required).
* Promote a blameless post\-mortem culture and continuous learning.
* Establish healthy and sustainable on\-call routines for the team.
**Automation and Platform Engineering**
* Drive adoption of IaC (Terraform, Ansible) and CI/CD pipelines.
* Reduce toil through automation and building reusable platform capabilities.
* Establish observability standards (logs, metrics, traces) using Prometheus, Grafana, Datadog, and OpenTelemetry.
Work Model: On-site \- Recife \- PE