Senior Reliability and Platform Engineer

Indeed

Company

Job typeFull-time
Workplace typeOnsite
Experience levelNo experience limit
Education levelNo degree limit

Description

Job Summary: We are seeking a Senior Reliability and Platform Engineer to enhance service and platform reliability, transforming incidents into structural improvements and engineering solutions. Key Highlights: 1. Driving the evolution of reliability and technological platform 2. Transforming incidents into structural improvements and automations 3. Focusing on engineering solutions for operational problems We seek a Senior Reliability and Platform Engineer to drive the reliability evolution of our services and our technological platform. The role involves transforming recurring incidents, alerts, and operational issues into structural improvements, automations, and preventive standards—not merely an operational profile focused solely on incident response, but rather an engineer who delivers engineering solutions. **Key Responsibilities** * Review and evolve alerts, monitors, and triggering criteria. * Ensure appropriate escalation and reduce ignored or unassigned alerts. * Implement and monitor SLIs, SLOs, and availability metrics. * Improve platform observability: logs, metrics, tracing, and APM. * Automate incident responses, diagnostics, and recoveries. * Create and maintain operational runbooks and playbooks. * Execute technical actions arising from post-mortems. * Reduce recurring failures and manual operational work (toil). * Support capacity, performance, resilience, and disaster recovery. * Evolve internal platform components and standards. * Provide technical support to DevOps, Platform, and Development teams. * Explore and apply AIOps solutions for anomaly detection and intelligent alert correlation. Requirements: * Senior experience in SRE, DevOps, or Platform Engineering. * Proficiency with Datadog or equivalent observability tools. * Experience with cloud platforms, Kubernetes, and infrastructure-as-code. * Knowledge of CI/CD, automation, scripting, or development. * Experience with incident management and root cause analysis. * Understanding of APIs, gateways, rate limiting, and distributed observability. * Ability to develop internal automations and tools. **Preferred Qualifications** * Practical experience with AIOps: anomaly detection, automatic alert correlation, or AI-assisted root cause analysis. * Experience in high-volume financial or retail environments. * Cloud certifications (GCP, AWS, or Azure) or Kubernetes certifications (CKA/CKAD). **Behavioral Profile** * Preventive and data-driven approach. * Ability to investigate and resolve complex problems. * Strong communication skills with both technical and business teams. * Autonomy to lead improvements without direct managerial oversight. * Systemic thinking and focus on risk reduction. Benefits Bradesco Health Insurance (co-payment) Odontoseg Dental Insurance (opt-in) Partnership with Dental Clinic ️ Meal Voucher or Food Voucher ️ Life Insurance PPR (Profit and Results Sharing) Commuter Benefit Bicycle Parking with Changing Room ️ 10% discount on Pernambucanas products TotalPass “Mãe Pernambucanas” Program with prenatal care support Partnerships with SESC and educational institutions for undergraduate and graduate courses Ângela Social Assistance Program supporting women experiencing violence.

Some content was automatically translated

Posted by

João Silva

Indeed · HR

Location

Similar jobs

Senior Reliability and Platform Engineer by Indeed in 2026 | ok.com