BtrBlocks: Efficient Columnar Compression for Data Lakes
Maximilian Kuschewski, Technical University of Munich
Analytics is moving to the cloud and data is moving into data lakes. These reside on blob storage services like S3 and enable seamless data sharing and system interoperability. To support this, many systems build on open storage formats like Apache Parquet. However, these formats are not optimized for remotely-accessed data lakes and today’s high-throughput networks. Inefficient decompression makes scans CPU-bound and thus increases query time and cost. With this work, we present BtrBlocks, an open columnar storage format designed for data lakes. BtrBlocks uses a set of lightweight encoding schemes, achieving fast and efficient decompression and high compression ratios.
About the speaker
Maximilian Kuschewski is a third-year Ph.D. student working with Prof. Viktor Leis at the Technical University of Munich. His research areas include efficient query processing, modern NVMe SSDs and cloud-native data analytics. Maximilian received his M.Sc. in Software Engineering from TUM, LMU and the University of Augsburg in 2020, and his B.Sc. in Computer Science and Engineering from the University of Augsburg in 2018.
Date & Time
Friday, November 3, 2023 - 14:00