Description
If you're passionate about solving problems and building solutions, and want to work with cutting-edge technologies in a dynamic environment, this role is for you!
We are a scale-up delivering outstanding results, growing rapidly while still having much to structure, create, test, adapt, and expand — a great opportunity to apply your knowledge, earn significant recognition, leave your mark on our story (and make a strong impression on your resume). We aim for ambitious outcomes in the coming years, and achieving them depends solely on us — working together and closely.
We are a company of innovative, creative, and diverse people, so we warmly welcome you — with all your uniqueness.
**What we value here:**
➡️ Embracing change;
➡️ Results orientation;
➡️ Closeness and collaboration;
➡️ Creating fans;
➡️ Conciseness and clarity;
➡️ Personal accountability.
**Your day-to-day responsibilities will include:**
* Monitoring alert channels (Google Chat, Dynatrace, Grafana, Kibana, and others), identifying anomalies, validating their legitimacy, and escalating to responsible parties or taking corrective actions;
* Using observability tools (Dynatrace, Grafana, Kibana, OpenSearch) to investigate logs, metrics, and traces in search of root causes;
* Providing **technical support to development teams**, investigating bugs and unexpected behaviors in production and staging environments;
* Participating in **war rooms** and incidents, supporting real-time triage, communication, and documentation;
* Overseeing deployments and validating service health post-release, identifying regressions or degradations;
* Collaborating with development teams on issues involving both application and infrastructure layers;
* Creating and maintaining **runbooks, checklists, and incident documentation**, contributing to the team’s knowledge base;
* Supporting the design and refinement of dashboards and alerts to improve observability coverage.
### **Our technology stack**
* **Cloud:** AWS (100% Cloud)
* **Orchestration:** Kubernetes, Docker
* **Observability:** Dynatrace, Grafana, Kibana, OpenSearch, Prometheus, CloudWatch
* **CI/CD:** GitHub Actions (preferred), Jenkins, Bitbucket CI
* **Version control:** GitHub (preferred), Bitbucket
* **Containerization:** Docker
* **Databases:** PostgreSQL, MySQL, RDS
* **Languages:** Python, PHP, Java, Bash
* **Messaging:** Kafka, SQS
### **✔️ What do you need to know?**
**Observability & Troubleshooting**
* Ability to read and interpret logs, metrics, and traces — understanding what the data implies even without explicit guidance;
* Basic or intermediate experience with tools such as Dynatrace, Grafana, Kibana, OpenSearch, CloudWatch, or similar;
* Capacity to correlate events: recognizing that a latency spike, a 500 error, and a deployment at 2 PM may all be part of the same incident;
**Development & Automation**
* Proficiency in **Python or Bash** for automation scripts, data analysis, or investigative support;
* Familiarity with **Git** (commits, branches, PRs) — you’ll be collaborating closely with developers;
* Basic understanding of how web applications work: REST APIs, HTTP/HTTPS, status codes, authentication;
* Willingness to read code even without being a developer — understanding an application’s flow greatly aids bug investigation;
**Infrastructure**
* Familiarity with **Linux** and daily command-line usage;
* Understanding of **Docker** and containers;
* Basic networking knowledge: DNS, load balancing, HTTP/HTTPS protocols;
* AWS experience is a plus — not required at expert level, but comfort navigating the console is helpful;
**Soft Skills (just as important as technical skills)**
* Investigative mindset: you won’t rest until you understand why a problem occurred;
* Strong communication: ability to clearly explain what’s happening, especially under pressure;
* Proactivity: you monitor, detect, and act — you don’t wait to be called;
* Ownership mentality: the systems are yours too.
### **➕ Nice-to-have qualifications (not mandatory, but they count)**
* Prior experience providing technical support to development teams or in production environments;
* Basic knowledge of Kubernetes;
* Familiarity with CI/CD tools;
* Experience with messaging systems (Kafka, SQS, RabbitMQ);
* Prior participation in war rooms or on-call shifts.