Description
**At Premiersoft, we turn challenges into real solutions!**
With over a decade of experience in development, we are driven by a clear purpose: **creating technological experiences that propel businesses** and accelerate our clients’ transformation.
Our team—comprising over 200 **\#Heroes**—combines technical excellence with our core DNA: **Team Player, Growth Driven, and Problem Solver**. We thrive on challenges, are guided by innovation, and are committed to the **delivery of high-impact solutions**, every single day.
**About the opportunity:**
You will work within one of the **largest cooperative ecosystems in the country**, in a highly regulated, data-driven environment where information is central to decision-making and the **evolution of financial products**.
The environment involves **large-scale data volumes, multiple integration sources, and real-world challenges around scalability, governance, and performance**. Your work will be directly tied to the **design and evolution of a modern data platform, focused on reliability, traceability, and availability for analytical and operational consumption.**
**Your responsibilities will include:**
* Developing and evolving **large-scale data pipelines with PySpark**, ensuring performance and reliability;
* Working with **distributed processing in Apache Spark and Databricks**, handling high-volume data;
* Integrating **multiple data sources (APIs, databases, systems)** with emphasis on consistency and traceability;
* Building and sustaining **Data Lake and Lakehouse architectures**;
* Implementing and optimizing **ETL/ELT processes** for analytical consumption;
* Ensuring **data quality, governance, and availability** across pipelines;
* Automating **monitoring, validation, and failure handling**;
* Optimizing **performance, scalability, and cost** of data routines.
**What you need:**
* Solid experience as a **Data Engineer** in high-volume environments;
* Strong expertise in **PySpark** for pipeline development and distributed processing;
* Production experience with **Apache Spark and Databricks**;
* Knowledge of **Data Lake and Lakehouse architectures**;
* Proficiency in **advanced SQL** for data transformation and validation;
* Experience with **pipeline orchestration (Airflow, ADF, or similar)**;
* Experience with **data modeling for analytical consumption**;
* Knowledge of **data quality, governance, and LGPD**;
* Hands-on experience with **version control (Git) and CI/CD applied to data**.
**Nice-to-have qualifications:**
* Experience with **financial sector data** (regulated, high-criticality environments);
* Experience with **real-time or near real-time processing** (Kafka, Spark Streaming);
* Experience in **high-volume, multi-integration environments**;
* Knowledge of **observability and monitoring of data pipelines**.
Remote hiring.
**What we offer:**
* A **collaborative** environment with **continuous knowledge sharing**;
* An open culture embracing **innovation, ideas, and ownership**;
* Use of **cutting-edge technologies** and industry best practices;
* Focus on **technical excellence** and **real-world impact** in deliverables;
* Ongoing encouragement of **learning and professional development**.
**Our benefits:**
* **TotalPass:** access to gyms, studios, and wellness activities;
* Health plan covering **mental health** services;
* **Paid time off:** 10 business days to recharge and take care of yourself;
* Gifts via Flash — **birthday gift** and **anniversary gift**;
* Referral bonus — **R$ 2.000,00 per hire**;
* Continuous development through **IDPs (Individual Development Plans), feedback, and certification support**;
* **Free English classes:** preparing you for international opportunities;
See what it’s like to join the Premiersoft team
Learn more about us
Visit our headquarters
Communication throughout the selection process occurs via email or WhatsApp. To avoid missing any updates, please add the domain **@premiersoft.net** to your list of trusted senders and monitor both your inbox and spam folder.