Description
Senior Data Engineering professional responsible for designing, developing, and optimizing scalable data pipelines for ingestion, transformation, and enrichment of structured and unstructured data, with a focus on supporting Artificial Intelligence solutions—especially Retrieval\-Augmented Generation (RAG) architectures and corporate knowledge base construction.
The role aligns with contractual Data Engineering responsibilities, applying techniques in advanced AI scenarios.
**Activities and Responsibilities**
Develop and implement data ingestion pipelines from databases, APIs, logs, and corporate document repositories (PDFs, HTML, textual documents).
Perform advanced data cleaning, transformation, enrichment, and versioning processes, ensuring integrity, traceability, and quality.
Design and maintain distributed pipelines using Apache Spark / PySpark on the Databricks platform and scalable data architecture (Data Lake / Lakehouse).
Implement AI-oriented data preparation strategies, including document segmentation (chunking), semantic enrichment, and integration with search and indexing mechanisms.
Support Data Science and AI/ML teams in preparing datasets for analytical and generative models.
Monitor and optimize performance, volume, and efficiency of data processing workflows.
Ensure compliance with best practices for data governance, retention, updating, and reliability.
**Mandatory Technical Knowledge**
Solid experience in Data Engineering
Python and/or PySpark
Apache Spark (batch and/or streaming)
Experience with ETL/ELT pipelines
Data modeling in Data Lake / Lakehouse environments
Experience consuming and integrating APIs
Experience in Cloud Computing environments (preferably Azure)
Use of version control (Git)
Desirable Knowledge (Technical Differentiators)
Experience with unstructured data (text and documents)
Experience with data pipelines for Artificial Intelligence
Knowledge of information retrieval strategies (RAG)
Integration with search and semantic indexing mechanisms
Experience with generative AI platforms (OpenAI, Azure OpenAI, or equivalents)
**Seniority Level**
Senior profile, capable of:
Defining data pipeline architecture
Proposing performance and quality improvements
Working with technical autonomy
Supporting and mentoring other data professionals
**Certifications**
Submit at least one (1) certification required by the contract, duly documented in the resume.
* Certified Data Management Professional (CDMP);
* Cloudera Certified Data Engineer (CCDE);
* AWS Certified Big Data;
* Microsoft Certified \- Azure Data Engineer Associate.
### **Employment Type:**
Contractor (PJ)
### **Department:**
Government