Asset Pipelines: A Modern Approach to Data Migration at Wormatia Worms

Asset Pipelines: A Modern Approach to Data Migration at Wormatia Worms

As an IT specialist and open-source enthusiast, I am passionate about innovative solutions that really make a difference. Today, I want to tell you about an exciting project in which we were able to unleash our passion for efficient data processing and migration with Dagster.

The Challenge: Migrating Legacy Data

Wormatia Worms, a football club with a long tradition, was facing a significant transformation. The old admin tool infrastructure needed to be updated and its data transferred to a modern system: Pimcore. The goal was to ensure a reliable, repeatable, and continuous migration that would meet requirements for up-to-date and accurate data.
The old admin tool had been in use for many years, and the club's archivist Christian Bub had entered many thousands of records into a system that, by today's standards, was no longer particularly modern.
The data was now to be transferred to a modern system to preserve it for the future and make it more accessible.
Ideally, the migration would also improve data quality and bring the data into a uniform format. Under no circumstances should important
data be lost, as the Wormatia archive is an important source for the club's history and is among the best of its kind in Germany.

The Innovative Approach: The Dagster Asset Pipeline

To enable the data migration, we chose Dagster. An asset pipeline seemed the ideal approach for achieving the necessary control and transparency. The data's journey began by retrieving information from the legacy API through a REST interface. Dagster then took over and transformed the data.

We also took the data dependencies into account in the pipeline. For example, club data and season information had to be imported before the player data could be imported. Altogether, there was a whole tree of dependencies. Dagster helped us import the data in the correct order and visualize this as well.

Dagster Asset Pipeline

Data quality was another problem, as the old system's database was not always consistent. In some cases, the data encodings were inconsistent too. Dagster's capabilities helped us clean and validate the data. The data had been stored in various systems and MySQL versions over what must have been 15 years.

In the past, people also liked using PHP hacks such as multiple utf8_decode / utf8_encode functions. We cleaned up this kind of data in the pipeline too.

Data Validation and Transformation

The data was not just transformed in the Dagster asset pipeline. A critical part of the process was validating data integrity – a crucial aspect of migrations that is easily overlooked but is enormously important for error-free data structures.

Because Dagster materializes the data, we were able to check data quality at every step and ensure the data was correct and consistent. This was particularly important because the data came from different sources and was being brought together in a new system. The great thing about the concept of data materialization is that we did not always have to retrieve the data from the legacy API, yet could still add new checks (including retrospectively). This allowed us, for example, to detect and clean up duplicates. We were also able to gradually clean up incorrectly encoded data this way.

Data validation in Dagster

We also used external libraries to transform the data. For example, we brought the data into a uniform format by using Pandas wherever possible. Pandas is a Python library that provides data structures and analysis tools, but also enables simple transformations.

After transformation and assessment in the asset pipeline, the data was sent to Pimcore through a GraphQL interface. Here, the focus was on continuous, maintainable delivery. Thanks to our pipeline system, we were able to update the process not only step by step but also continuously in the background, for example through a cron job.

Data in Pimcore

But even the best process is only as good as the data it processes. That is why we also checked data quality in Pimcore and made sure the data was correct and consistent there as well.
Christian Bub was a great help here, as he knows the data very well and gave me excellent support with validation and transformation.

WordPress Integration

Once the data was available in Pimcore, the next step was integrating it into the Wormatia Worms website, which is based on WordPress. To do this, we developed a plugin that incorporates the GraphQL data directly – a real highlight that combines the flexibility of WordPress with the power of Pimcore. The data can be integrated directly through a WordPress shortcode or visually through a Gutenberg element.

Another goal was to extend the GraphQL interface so that it provides pre-aggregated statistics and data. This significantly improved efficiency and also made it possible to connect the Wormatia live ticker, an in-house development, seamlessly to Pimcore.

Wormatia archive example
GraphQL query example

Conclusion: An Iterative and Scalable Migration Process

This project showed how using modern technology allows us not only to improve existing processes but to create entirely new possibilities for data integration. The combination of Dagster and Pimcore ensures that Wormatia Worms is ready for the future, while also providing the innovative strength needed for lasting success.

Do you need similar solutions or have questions about Dagster and integrated platforms? Let us know – we are happy to share our experiences!