Description
Job Summary:
You will be responsible for implementing, configuring, and administering observability solutions using Datadog, monitoring critical environments, and proposing continuous improvements.
Key Highlights:
1. Implementation and management of observability solutions with Datadog
2. Creation of dashboards and proactive monitoring of critical environments
3. Performance analysis, troubleshooting, and observability automation
Permanent Regular
Job Description:
Pluxee is a global player in employee benefits and engagement that operates in 31 countries. Pluxee helps companies attract, engage, and retain talent thanks to a broad range of solutions across Meal \& Food, Wellbeing, Lifestyle, Reward \& Recognition, and Public Benefits.
Powered by leading technology and more than 5,000 engaged team members, Pluxee acts as a trusted partner within a highly interconnected B2B2C ecosystem made up of more than 500,000 clients, 36 million consumers and 1\.7 million merchants.
Conducting its business as a trusted partner for more than 45 years, Pluxee is committed to creating a positive impact on all its stakeholders, from driving business to local communities, to supporting wellbeing at work for employees while protecting the planet.
You will be responsible for implementing, configuring, and administering observability solutions using Datadog.
**Responsibilities and Duties:**
* Create and maintain executive, operational, and technical dashboards for monitoring critical environments;
* Develop and fine-tune intelligent alerts, reducing false positives and increasing operational effectiveness;
* Monitor applications, infrastructure, and services in cloud and hybrid environments;
* Analyze logs, metrics, and traces for troubleshooting and identification of performance bottlenecks;
* Support development, operations, and SRE teams in investigating and resolving critical incidents;
* Implement observability in Kubernetes environments, Docker containers, and microservices architectures;
* Automate observability and provisioning processes using tools such as Terraform and Ansible;
* Conduct root cause analyses (RCA) and support post\-mortem processes;
* Propose continuous improvements in availability, reliability, and user experience;
* Integrate Datadog with corporate tools, CI/CD pipelines, and cloud platforms;
* Define and track availability indicators, SLIs, SLOs, and SLAs;
* Participate in assisted operations, war rooms, and monitoring of critical business events;
* Document processes, monitoring standards, and observability best practices.
**Key Competencies:**
* Datadog implementation and management (APM, Logs, Metrics, RUM, Synthetic Monitoring);
* Dashboard creation and proactive monitoring;
* Troubleshooting and performance analysis;
* Integration with AWS, Azure, GCP;
* Kubernetes and Docker;
* CI/CD and automation (Terraform, Ansible, etc.);
* SRE and observability practices;
* Completed undergraduate degree in Computer Science, Computer Engineering, Software Engineering, Information Systems, Systems Analysis and Development, or related fields.
* Technical English: desirable.
**Preferred Qualifications:**
* Cloud certifications (AWS, Azure, GCP)
* Experience with complementary tools (Prometheus, Grafana, ELK Stack)
* Knowledge of security and compliance monitoring
**Additional Information**
--------------------------
Our benefits are a market differentiator:
* Meal Pass
* Food Pass
* Medical and dental assistance
* Life insurance
* Culture Pass
* Wellhub
* Support Pass
* Transportation Allowance (VT)
* 50% subsidy for medication costs