Parquet files: The powerful format for data lakes

Parquet files: The powerful format for data lakes

People like to say these days that data is the oil of the future. In fact, data has always been important in IT. That is why it is important to take a close look at data. As data volumes become larger and keep growing, the way we store and process data becomes increasingly crucial. One data format that proves extremely useful here is Parquet. The data format is also frequently used by data lakes. Let's take a closer look at the format.

What are Parquet files?

Parquet is a column-based storage format originally developed by Twitter (now X) and later continued as an open-source project under the Apache license. It is specifically designed to store large volumes of data efficiently and provide optimal performance when accessing data (read access). Traditional row-based formats store data vertically in rows. Parquet organizes data in a hybrid format. It is a mixture of horizontal and vertical storage. This makes it particularly suitable for massive data analyses and fast queries. The data is partitioned horizontally and vertically.

Example: Imagine a huge table in which each record contains information about different customers. If you only want to query the "customer number" column, the row-based format has to go through every row, while Parquet can load just the "customer number" column. This saves time and computing resources.

Let's go into partitioning in a little more detail.

Partitioning

Suppose we have a CSV file with customer data. It could look like this very simple example:

id,name,region,land,kundennummer
1,Alice,EMEA,DE,10000
2,Bob,EMEA,FR,10001
3,Charlie,APAC,JP,10002
4,Diana,EMEA,DE,10003
5,Eric,APAC,JP,10004

With Parquet, it could then be stored like this:

kunden_parquet/
├── region=EMEA/
│   ├── land=DE/
│   │   ├── part-00000-...-c000.parquet
│   │   └── part-00001-...-c000.parquet
│   └── land=FR/
│       ├── part-00000-...-c000.parquet
│       └── ...
└── region=APAC/
    └── land=JP/
        ├── part-00000-...-c000.parquet
        └── ...

You can define how the data is partitioned yourself. In this example, it was partitioned by "region" and then by "land".

The structure of Parquet files

Parquet uses several techniques to optimize storage capacity and maximize read and write speeds. Here are the most important techniques:

  1. Column-based storage: Data is stored by column, meaning all values of a particular column are stored together. This enables better compression and faster read times, particularly for analytical queries.

  2. Compression: Parquet supports several compression techniques (such as Snappy, Gzip, or LZ4), which let you minimize the amount of storage needed. Good compression can provide significant savings for large datasets.

  3. Schema: Parquet uses a very flexible schema that allows complex data types such as nested structures or arrays to be stored. The schema is stored in the file header, so the structure is available at any time.

  4. Metadata: Parquet stores important metadata that also helps access the data more efficiently. This includes information about the number of rows, compression, and the schema, which considerably simplifies data analysis.

Use cases for Parquet files

Parquet has established itself as the preferred format for data lakes and data-intensive applications. There is a wide range of possible uses:

  • Big data analytics: With its column-based storage and excellent compression, Parquet enables efficient data processing with tools such as Apache Spark or Apache Hive. If you want to analyze large volumes of data, Parquet is often the first choice. It is almost a kind of industry standard.

  • Data warehousing: In data warehouses, where large volumes of data are retrieved along various dimensions, Parquet offers the necessary speed and efficiency.

  • Machine learning: When processing training data for machine learning models, Parquet can help store data optimally and access it quickly.

  • Data integration: Many modern data integration tools support Parquet directly. This considerably simplifies its use in ETL processes.

Tools

If you have data in Parquet format and want to take a look inside, you can use one of the following tools, for example.
Here are a few tools I really enjoy using.

parquet-cli

The tool can display the schema of a Parquet file (or directory). The data can then be output in CSV format.
Data can be queried and filtered.

https://github.com/chhantyal/parquet-cli

DuckDB

DuckDB is what is known as an in-process SQL engine. I enjoy using DuckDB in Jupyter notebooks from time to time, because it makes it very easy to use SQL to access data in different formats (including CSV). With DuckDB, you can also access Parquet files and then select the columns directly using SQL. Very handy.

https://duckdb.org/

PyArrow library

In the Python world, the numpy and pandas libraries are very well known when it comes to data analysis.

The pyarrow library (also from Apache) can be a good addition here.

Here is an example where we create a Pandas DataFrame, which is very common.

import numpy as np
import pandas as pd
import pyarrow as pa

df = pd.DataFrame({'one': [-1, np.nan, 2.5],
                   'two': ['foo', 'bar', 'baz'],
                   'three': [True, False, True]},
                   index=list('abc'))

Using pyarrow, we can now very easily write a Pandas DataFrame to a Parquet file.

import pyarrow.parquet as pq

table = pa.Table.from_pandas(df)
pq.write_table(table, 'example.parquet')

https://arrow.apache.org/


If you would like more information and want to understand the exact structure of the format, you can find more details on the official website at https://parquet.apache.org/.