Faster chat, better deals — Get the App

Senior SRE Analyst | Remote

Indeed

Company

Job typeFull-time
Workplace typeRemote
Experience level1 to 2 years
Education levelBachelor's Degree

Description

If you're passionate about solving problems and building solutions, and want to work with cutting-edge technologies in a dynamic environment, this role is for you! We are a scale-up delivering outstanding results, growing rapidly while still having much to structure, create, test, adapt, and expand — a great opportunity to apply your knowledge, earn significant recognition, leave your mark on our story (and make a strong impression on your resume). We aim for ambitious outcomes in the coming years, and achieving them depends solely on us — working together and closely. We are a company of innovative, creative, and diverse people, so we warmly welcome you — with all your uniqueness. **What we value here:** ➡️ Embracing change; ➡️ Results orientation; ➡️ Closeness and collaboration; ➡️ Creating fans; ➡️ Conciseness and clarity; ➡️ Personal accountability. **Your day-to-day responsibilities will include:** * Monitoring alert channels (Google Chat, Dynatrace, Grafana, Kibana, and others), identifying anomalies, validating their legitimacy, and escalating to responsible parties or taking corrective actions; * Using observability tools (Dynatrace, Grafana, Kibana, OpenSearch) to investigate logs, metrics, and traces in search of root causes; * Providing **technical support to development teams**, investigating bugs and unexpected behaviors in production and staging environments; * Participating in **war rooms** and incidents, supporting real-time triage, communication, and documentation; * Overseeing deployments and validating service health post-release, identifying regressions or degradations; * Collaborating with development teams on issues involving both application and infrastructure layers; * Creating and maintaining **runbooks, checklists, and incident documentation**, contributing to the team’s knowledge base; * Supporting the design and refinement of dashboards and alerts to improve observability coverage. ### **Our technology stack** * **Cloud:** AWS (100% Cloud) * **Orchestration:** Kubernetes, Docker * **Observability:** Dynatrace, Grafana, Kibana, OpenSearch, Prometheus, CloudWatch * **CI/CD:** GitHub Actions (preferred), Jenkins, Bitbucket CI * **Version control:** GitHub (preferred), Bitbucket * **Containerization:** Docker * **Databases:** PostgreSQL, MySQL, RDS * **Languages:** Python, PHP, Java, Bash * **Messaging:** Kafka, SQS ### **✔️ What do you need to know?** **Observability & Troubleshooting** * Ability to read and interpret logs, metrics, and traces — understanding what the data implies even without explicit guidance; * Basic or intermediate experience with tools such as Dynatrace, Grafana, Kibana, OpenSearch, CloudWatch, or similar; * Capacity to correlate events: recognizing that a latency spike, a 500 error, and a deployment at 2 PM may all be part of the same incident; **Development & Automation** * Proficiency in **Python or Bash** for automation scripts, data analysis, or investigative support; * Familiarity with **Git** (commits, branches, PRs) — you’ll be collaborating closely with developers; * Basic understanding of how web applications work: REST APIs, HTTP/HTTPS, status codes, authentication; * Willingness to read code even without being a developer — understanding an application’s flow greatly aids bug investigation; **Infrastructure** * Familiarity with **Linux** and daily command-line usage; * Understanding of **Docker** and containers; * Basic networking knowledge: DNS, load balancing, HTTP/HTTPS protocols; * AWS experience is a plus — not required at expert level, but comfort navigating the console is helpful; **Soft Skills (just as important as technical skills)** * Investigative mindset: you won’t rest until you understand why a problem occurred; * Strong communication: ability to clearly explain what’s happening, especially under pressure; * Proactivity: you monitor, detect, and act — you don’t wait to be called; * Ownership mentality: the systems are yours too. ### **➕ Nice-to-have qualifications (not mandatory, but they count)** * Prior experience providing technical support to development teams or in production environments; * Basic knowledge of Kubernetes; * Familiarity with CI/CD tools; * Experience with messaging systems (Kafka, SQS, RabbitMQ); * Prior participation in war rooms or on-call shifts.

Some content was automatically translated

Posted by

João Silva

Indeed · HR

Location

João Silva

Indeed · HR

Similar jobs

Senior SRE Analyst | Remote job by Indeed in 2026 | ok.com