Faster chat, better deals — Get the App

Senior SRE (Afternoon/Night Shift)

Indeed

Company

Job typeFull-time
Workplace typeHybrid
Experience levelMore than 10 years
Education levelBachelor's Degree

Description

At Banco ABC Brasil, we believe in the authenticity of each individual. After all, we have our own way of doing things, of relating to others, of transforming businesses and building a sustainable future—in an inclusive, respectful, and welcoming manner. Because we genuinely care about people and build authentic relationships grounded in trust and closeness. If you are passionate about challenges and seek an environment where you can grow professionally—with autonomy to lead major projects and take ownership of your career—this is the place for you! With us, you’ll have daily opportunities to collaborate with financial industry specialists and receive guidance and support from strategic leaders to shape your future and contribute jointly to our growth. **We believe that caring for our employees is the key to success. That’s why we offer:** * Benefits that truly make a difference * Development opportunities * An inspiring work environment We are seeking a **Senior SRE** with a hands-on profile to assume frontline responsibility for the reliability and stability of our most critical systems. In this role, you will be the **guardian and technical investigator** of our ecosystem. You will resolve highly complex incidents in multi\-cloud environments, playing a **vital role in advanced troubleshooting across our infrastructure**. Your responsibility is to master low-level operations—ensuring efficiency, security, and availability aligned with SRE culture. **Working hours: 2:00 PM to 11:00 PM.** Are you ready to join a team that transforms challenges into opportunities? Come with us! **Responsibilities and Duties** Capacity, Performance, and Availability Management * Continuously plan and adjust computing resource capacity (compute, memory, storage, network) on AWS and on\-premises—anticipating bottlenecks and avoiding waste. * Monitor, analyze, and optimize service and infrastructure performance—identifying degradations before they impact end users (applying USE and RED methodologies). * Define, implement, and maintain SLOs, SLAs, and error budgets—ensuring high availability through automation and well-documented runbooks. * Build and maintain automated controls that guarantee agreed-upon reliability KPIs with the business—including full traceability and auditability. Automation and Reliability Engineering * Develop and evolve operational automations—reactive and predictive scaling, automatic remediation, zero\-touch provisioning—to reduce toil and increase resilience. * Manage and optimize EKS clusters: provisioning, scalability (HPA / VPA / Cluster Autoscaler / Karpenter), networking, storage, and production workload troubleshooting. * Ensure versioned, reproducible, and auditable infrastructure. * Conduct chaos engineering to validate system resilience (controlled failure simulations, game days). Observability * Maintain complete observability stacks: metrics, logs, distributed tracing, and SLO\-oriented alerts. * Build dashboards and alerts using Prometheus, Grafana, and CloudWatch—with end\-to\-end visibility into infrastructure health. FinOps and Cost Management * Apply cloud cost optimization concepts and practices: rightsizing, reserved instances, savings plans, and spot instances. * Generate cost reduction reports and recommendations for AWS—using AWS Cost Explorer, Kubecost, or equivalent tools. * Implement tagging and chargeback mechanisms to enable cost visibility per service, squad, or product—fostering FinOps culture within the team. Incident Response and Technical Leadership * Participate in on\-call rotations, lead resolution of high\-severity incidents, and conduct blameless post\-mortems with concrete action items. * Support fellow SREs by disseminating reliability, observability, and operational engineering practices. * Serve as a technical reference for infrastructure architecture decisions related to reliability, capacity, and performance. * Collaborate with the cloud engineering team on infrastructure technical reviews. **Requirements and Qualifications** **Expected Technology Stack / Tools** **Core** **AWS Cloud:** EC2, Auto Scaling, EKS, Lambda, RDS/Aurora, S3 (lifecycle/tiers), EBS (gp3/io2\), EFS/FSx, VPC, Transit Gateway, ALB/NLB, Route53, IAM/SCP, CloudWatch, AWS Backup **Kubernetes / EKS:** EKS, Helm, Kustomize, HPA, VPA, Cluster Autoscaler, Karpenter, Network Policies, CSI Drivers, Persistent Volumes, Istio or Linkerd (desirable) **Storage — Cloud and On\-premises:** EBS (gp3/io2\), EFS, FSx, S3 lifecycle, CSI Drivers, SAN/NAS/NFS on\-premises, Ceph (desirable), AWS Backup, Commvault. **Infrastructure as Code:** Terraform, Ansible, CloudFormation **CI/CD and GitOps:** GitHub Actions, Azure DevOps, ArgoCD, Flux **Observability:** Prometheus, Grafana, Dynatrace or Datadog, CloudWatch, CloudTrail. **FinOps:** AWS Cost Explorer, Rightsizing, Reserved Instances, Savings Plans, Spot Instances. **On\-Premises:** VMware vSphere/ESXi, Bare\-metal Linux (Ubuntu, RHEL), Corporate Networks (VLAN, basic BGP/OSPF), Dell EMC / HPE (desirable) **Languages / Scripting:** Python, Bash/Shell. **Security (SRE\-scope):** IAM/SCP, RBAC in Kubernetes, Secrets Manager, Parameter Store, network policies. **Technical Competencies** * Solid experience managing capacity and performance in hybrid environments (cloud \+ on\-premises), with proven accountability for SLOs and KPIs. * Advanced expertise in AWS: compute, storage, networking, IAM, and managed services at production scale. * Production-grade Kubernetes/EKS: provisioning, troubleshooting, scaling, and storage—with minimum 4 years’ experience. * Production-level Terraform: modules, remote state, workspaces, and drift reconciliation. * End\-to\-end observability: metrics, logs, tracing, SLO\-oriented alerts, and operational dashboard creation. * Hybrid storage: deep knowledge of EBS, EFS, and FSx in cloud—and SAN/NAS/NFS on\-premises—including IOPS and capacity planning. * Python or Bash for automation and internal tooling. **Differentiators** * Multi\-cloud experience (AWS \+ Azure or AWS \+ GCP). * Production knowledge of service mesh (Istio or Linkerd). * Experience with FinOps tools (Kubecost, CloudHealth, Spot.io). * Participation in open source communities or meaningful contributions on GitHub. * Experience with event\-driven architecture (Kafka/MSK, SQS/SNS) in an SRE context. **Soft Skills** * Data\-driven, analytical thinking focused on reliability metrics and KPIs. * Clear and concise communication with technical teams and business stakeholders. * Autonomy and proactivity in highly complex and ambiguous environments. * Technical leadership without formal authority—exerting influence through expertise. * Resilience and focus under pressure during critical incidents. * Collaborative mindset and genuine willingness to mentor and share knowledge. **Certifications** Candidates must hold at least one certification in either SRE or AWS Cloud domains. The complete absence of certifications in both areas—without a strong, demonstrable technical portfolio—is a disqualifying factor. AWS Solutions Architect (Associate or Professional) carries the strongest weight among cloud certifications. AWS Solutions Architect Associate or Professional \- **Strong Differentiator** AWS DevOps Engineer Professional \- Differentiator AWS SysOps Administrator Associate \- Differentiator Certified Kubernetes Administrator (CKA) \- Differentiator Certified Kubernetes Application Developer (CKAD) \- Differentiator HashiCorp Terraform Associate \- Differentiator **Academic Background** * Bachelor’s degree in Computer Science, Software Engineering, Network Engineering, or related fields. * Postgraduate degrees, MBAs, or recognized technical specializations are differentiators. **Additional Information** * Health Insurance; * Omint Dental Insurance; * Life Insurance; * Profit Sharing Plan (PLR); * Performance Bonus (PPR); * ABC com Você: a program supporting employees and their families with legal, social, psychological, and financial assistance; * Meal Allowance; * Food Allowance; * Extended Paternity and Maternity Leave: 20 days paternity leave and 6 months maternity leave; * Childcare/Babysitter Assistance; * Annual Day Off; * Home Office Infrastructure Support; * TotalPass; We are ABC Brasil—the multi\-service bank with over 35 years of history, specializing in financial solutions and empowering major Brazilian businesses—combining international solidity with the agility of local, close, and autonomous management. With a comprehensive portfolio of products and services, our focus is on delivering real impact to our customers—evolving alongside the market and adapting to each customer’s unique needs, always with responsibility, integrity, and mutual trust. This approach to relationships makes us unique. We believe authentic connections—rooted in respect for differences—create a collaborative, human, and inspiring environment. Here, every person can be themselves—and grow with autonomy and ownership. **ABC Brasil. The bank for those who are singular.** \#EuSouSingular \#SouABCBrasil \#ABCBrasil

Some content was automatically translated

Posted by

João Silva

Indeed · HR

Location

João Silva

Indeed · HR

Similar jobs

Senior SRE (Afternoon/Night Shift) job by Indeed in 2026 | ok.com