Building Data Products: The Data Products Series Volume 2, by Mario Meir-Huber
Create trustworthy, production-ready data that remains useful long after the first pipeline run. This practical book helps data leaders, architects, engineers, and product teams design and implement the retrieval foundation needed to deliver reliable information across an organization.
Introduction
Chapter 1: A Technical Definition of Data Products
Chapter 2: The G in Retrieval
Chapter 3: The A in Retrieval
Chapter 4: The P in Retrieval
Chapter 5: The Sample Project
Apply the GAP triad of Governance, Architecture, and People to the realities of inconsistent source systems, changing schemas, unclear ownership, regulatory requirements, and uneven data quality. Profile sources before connecting to them, classify information during ingestion, establish meaningful quality gates, automate metadata capture, document lineage, and manage retention and audit logging as part of delivery rather than as an afterthought.
Design a reusable data pipeline architecture that supports batch processing, APIs, and change data capture. Evaluate connector options, orchestration patterns, storage choices, quarantine processes, and schema validation. Instead of developing a different pipeline for every use case, configure a shared template that embeds consistent controls and scales without multiplying engineering effort.
Coordinate the people behind the technology. Map responsibilities with RACI, establish productive agreements with source system teams, involve domain experts in business metadata, and define how teams respond when pipelines or quality controls fail. A detailed sample project connects the concepts using realistic source systems, YAML definitions, and an open-source stack that readers can run and examine themselves.
This book replaces fragmented practices with an implementable approach to DataOps, metadata management, data lineage, quality assurance, and dependable ingestion. Build the governed retrieval capability your organization needs to put trusted information into action.
Mario Meir-Huber helps organizations turn scattered data initiatives into governed, business-driven Data Products. A former Head of Data and ex-Microsoftie, he has built Data Products across European companies. Mario lectures at WU Vienna, teaches on LinkedIn Learning, and speaks at events like WeAreDevelopers, Data Modeling Zone, London Tech Week, Data Science Conference, and GITEX. He co-authored The Handbook of Data Science and AI and is writing The Data Products Series. You can follow his Substack at https://futuregrade.substack.com for interesting insights on Data Products and AI.
Please complete all fields.