Faster chat, better deals — Get the App

IT Specialist (Observability Architecture/SRE)

Indeed

Company

Job typeFull-time
Workplace typeOnsite
Experience levelNo experience limit
Education levelNo degree limit

Description

Job Summary: Observability/SRE Specialist to define, implement, and evolve observability standards, manage Kubernetes clusters, and messaging platforms. Key Highlights: 1. Define, implement, and evolve observability standards 2. Manage and optimize Kubernetes clusters with Service Mesh 3. Administer and evolve messaging platforms (Kafka, RabbitMQ) IT Specialist (Observability Architecture/SRE) Country: Brazil **\# WHO WE ARE** F1RST is the future—and your career is here! Our culture is built on “People, Innovation, and Results.” We deliver services and experiences for over 60 million customers across the entire Santander ecosystem. Join the team whose purpose is to support people and help businesses thrive. We are passionate about technology. We are F1RST Digital Services. Follow our LinkedIn to stay updated on all news: https://www.linkedin.com/company/f1rstdigitalservices We have an opening for you to become an **Observability/SRE Specialist**. **Your role here will be:** * Define, implement, and evolve observability standards for applications and infrastructure * Deploy and maintain monitoring and tracing stacks (Prometheus, Grafana, Dynatrace, OTEL, Jaeger, ELK/Kibana) * Manage and optimize Kubernetes (EKS) clusters with Service Mesh (Istio), ensuring visibility and traffic control between services * Administer and evolve messaging platforms (Kafka and RabbitMQ), ensuring availability, performance, and reliability * Contribute to defining application instrumentation standards using OpenTelemetry * Monitor and analyze metrics, logs, and traces to anticipate incidents and drive continuous improvement * Perform troubleshooting in distributed environments, identifying performance bottlenecks and inter-service communication failures * Support development squads in adopting observability best practices and event-driven architecture * Automate provisioning and resource management via Operators on Kubernetes * Document architectures, technical standards, and operational procedures **Mandatory Requirements:** * Advanced observability expertise: implementation and maintenance of metrics, logs, and distributed tracing in cloud-native architectures * Experience with monitoring tools: Dynatrace, Grafana, Prometheus, and Kibana—including dashboard configuration, alerting, SLO/SLI definition, and troubleshooting * Knowledge of observability tools: OpenTelemetry (OTEL) for application instrumentation and Jaeger for distributed trace analysis * Solid Kubernetes (EKS) experience: cluster administration, troubleshooting, resource tuning, namespace management, and policy enforcement * Service Mesh knowledge: Istio—including traffic control, mTLS, security policies, service observability, and sidecar management * Experience with messaging platforms: Apache Kafka (Confluent), including use of tools such as Kafdrop for inspection and troubleshooting, and RabbitMQ (Amazon MQ) * Knowledge of Kubernetes Operators for automating deployment and management of stateful workloads (e.g., Kafka Operator, RabbitMQ Operator) * Experience with event-driven architecture (EDA) and microservices—including patterns such as pub/sub, consumer groups, DLQ, and retry strategies * Advanced distributed environment performance and latency troubleshooting skills * Incident management experience, root cause analysis (RCA), and preventive action plan definition **Desirable Requirements:** * Kubernetes certifications (CKA, CKAD) or observability specializations * Practical knowledge and usage of market AI tools to boost productivity—such as Devin, Claude, Cursor, ChatGPT Enterprise, GitHub Copilot, etc. * Experience with capacity management and performance tuning for Kafka and RabbitMQ clusters * Knowledge of high availability and disaster recovery strategies for messaging platforms * Experience integrating metrics and traces into CI/CD pipelines * Knowledge of Kubernetes and Service Mesh security (mTLS, RBAC, policies) * Experience with multi-cluster and multi-region environments * Advanced English and Spanish for collaboration with global teams **Work Location:** Geração Digital – Av Interlagos, 3501 – Interlagos, São Paulo \- SP **\# BENEFITS:** ➡️ Meal allowance; ➡️ Health insurance; ➡️ Dental insurance: Basic and intermediate plans; ➡️ Transportation allowance; ➡️ Flexible Vacation: 24 business days of vacation, divisible into up to 6 periods; after every 2 months worked, you may enjoy 4 business days; ➡️ Birthday Day Off; ➡️ Profit Sharing Program (PPR); ➡️ Gym partnerships: Wellhub, Totalpass; ➡️ Flexible Working: Hybrid work model—2 days remote and 3 days onsite; ➡️ Training platforms with over 100,000 courses; ➡️ Career paths for professional development; ➡️ Flex Learning: Exclusive study incentive program for High-Performance employees; ➡️ Childcare allowance; ➡️ “Nascer” Program and Extended Paternity Leave; ➡️ Life insurance; ➡️ “Nascer” Program; ➡️ Be Healthy—Program encouraging healthier habits; ➡️ PAPE—Specialized Personal Support Program; \#LI\-Hybrid

Some content was automatically translated

Posted by

João Silva

Indeed · HR

Location

João Silva

Indeed · HR

Similar jobs

IT Specialist (Observability Architecture/SRE) by Indeed in 2026 | ok.com