Description
Hybrid Opportunity\- Avenida Paulista\- 2X per Week
We are seeking a Data Tech Lead with experience in designing, architecting, and implementing the decisional layer on the **Databricks** platform. This professional will be responsible for guiding the data engineering strategy, ensuring technical excellence in modeling, multi\-source ingestion, and large\-scale processing, while also serving as the technical interface between the client’s architects and the execution team.
**Responsibilities and Duties:**
* Design and implement data architecture on the Databricks platform using **Medallion Architecture** patterns (Bronze, Silver, Gold).
* Define ingestion strategies for multiple sources (Teradata, DB2, SQL Server, Oracle, SAS, HIVE), ensuring integrity via **CDC** and **Streaming**.
* Establish data governance, observability, and integration contract standards via **REST API**.
* Provide technical leadership to mid\- and junior\-level engineers, conduct code reviews, and ensure adherence to engineering best practices.
* Serve as the technical liaison with architects and stakeholders to align solutions and requirements.
* Structure the **Context Store** and define official data sources by domain.
* Implement and monitor automated data quality tests (using **Great Expectations** or **DLT expectations**).
* Optimize workload cost and performance (cluster sizing and processing efficiency).
* Manage technical complexity of data ingestion from heterogeneous ecosystems (Mainframe, On\-premises, and Cloud).
* Ensure full automation of the data lifecycle through CI/CD pipelines.
**Requirements:**
* Proven experience with **Databricks** (Unity Catalog, DLT \- Delta Live Tables, Workflows, and Jobs).
* Advanced proficiency in **PySpark** (cluster optimization, broadcast joins, partitioning, and AQE).
* Advanced hands\-on experience with **SQL** (Window Functions, CTEs, Query Tuning).
* Deep expertise in **Data Modeling** (Dimensional, Data Vault, and Medallion).
* Practical experience with **CDC** tools (Debezium, GoldenGate, or Lakeflow Connect) and **Streaming** platforms (Kafka or Azure Event Hub).
* Familiarity with REST API integration (OAuth/JWT, Idempotency).
**Nice to Have:**
* **Databricks Certified Data Engineer Professional** certification.
* Experience migrating **Hive Metastore** to **Unity Catalog**.
* Knowledge of advanced performance tuning (**Z\-Ordering**, **Liquid Clustering**).
* Prior experience in Insurance or Banking sector projects.
**Technologies:**
* **Primary Platform:** Databricks (Delta Lake, Unity Catalog, DLT, Lakeflow).
* **Languages:** Python (PySpark) and SQL.
* **Cloud:** Azure (Event Hub, Storage).
* **Databases and Sources:** Teradata, DB2, Oracle, SQL Server, SAS, and Hive.
* **Streaming/Messaging:** Kafka and Structured Streaming.
* **Quality and Governance:** Great Expectations and Unity Catalog.