DUAL-STAGE CONTRASTIVE SELF-SUPERVISED SEMANTIC ETL FRAMEWORK FOR ADAPTIVE AND EFFICIENT DATA LAKEHOUSE PROCESSING

Main Article Content

K.Abirami, S. Punitha

Abstract

The rapid adoption of data lakehouse architectures has intensified the need for efficient Extract–Transform–Load (ETL) mechanisms capable of handling heterogeneous, large-scale data while supporting adaptive analytics. However, conventional ETL frameworks rely on static transformation rules and lack semantic awareness, leading to inefficient data integration and suboptimal query performance. Existing approaches also struggle to leverage unlabeled data for intelligent transformation, creating a gap in self-supervised and context-adaptive ETL design. To address these limitations, this study proposes a Dual-Stage Contrastive Self-Supervised Semantic ETL (DSC-SS-ETL) framework that integrates contrastive learning into both data and query processing stages. The first stage learns semantic embeddings of heterogeneous data through self-supervised contrastive representation learning, while the second stage aligns query patterns with data representations for adaptive transformation and optimized scheduling. The framework incorporates a semantic encoder, dual-stage contrastive module, adaptive transformation engine, and intelligent scheduler. Experiments were conducted using TPC-DS benchmark datasets and real-world cloud warehouse workloads under varying query conditions. The proposed model achieves 96.8% semantic alignment accuracy, 95.4% transformation precision, and a 33.7% reduction in ETL latency compared to baseline methods. These results demonstrate the effectiveness of contrastive self-supervised learning in enhancing ETL efficiency and adaptive analytics.

Article Details

Section
Articles