Description
Job Summary:
A professional responsible for monitoring systems, analyzing and resolving Level 1 incidents, and ensuring service availability and performance.
Key Highlights:
1. System monitoring and analysis
2. First-level incident response and resolution (Level 1)
3. SLA communication and tracking
**Job Responsibilities:**
* Monitor servers, applications, databases, networks, and critical services.
* Monitor dashboards and monitoring tools to identify alerts and anomalies.
* Provide first-level incident response and analysis (Level 1 or NOC).
* Log, classify, and escalate tickets to specialized teams when necessary.
* Perform initial diagnostic procedures to identify potential root causes of failures.
* Track incident resolution and ensure SLA compliance.
* Escalate critical issues according to company procedures.
* Monitor automated routines, backups, integrations, and scheduled processes.
* Prepare reports on availability, performance, and incidents.
* Update operational documentation and support procedures.
* Participate in preventive actions to reduce downtime and recurrence of failures.
* Validate service normalization following corrections or maintenance activities.
* Communicate service outages and incident status to internal or external customers.
**Requirements:**
* Problem analysis and resolution skills.
* Basic knowledge of networking, operating systems, and databases.
* Strong communication skills for incident handling.
* Organizational skills and attention to detail.
**Preferred Qualifications:**
* Practical knowledge and experience with Zabbix, Grafana, n8n, and other infrastructure and service monitoring and automation tools.
Benefits:
* Dental insurance
* Commercial partnerships and discounts
* Profit-sharing program
* Life insurance
* Meal allowance
* Food allowance
* Transportation allowance
Work Location: On-site