Building Data Products: The Data Products Series Volume 2

Original price was: $39.95.Current price is: $34.99.
$39.95

Building Data Products: The Data Products Series Volume 2, by Mario Meir-Huber

Create trustworthy, production-ready data that remains useful long after the first pipeline run. This practical book helps data leaders, architects, engineers, and product teams design and implement the retrieval foundation needed to deliver reliable information across an organization.

Topics

Introduction

  • Defining Data Products
  • The Data Product process
  • How to use this book

 

Chapter 1: A Technical Definition of  Data Products

  • Open Data Product Standard
  • Processes for Data Products
  • Common pitfalls in creating Data Products
  • Catalog-based updates of a Data Product
  • Pipeline-based updates of a Data Product
  • Key learnings

 

Chapter 2: The G in Retrieval

  • The starting point: source system profiling
  • The source system scorecard
  • Classifying data at ingestion
  • Data quality gates in data retrieval
  • Automating metadata capture
  • Lineage from source to landing zone
  • Retention, compliance, and audit logging
  • Key learnings

 

Chapter 3: The A in Retrieval

  • The internals of the retrieval pipeline
  • The reference architecture
  • Orchestration architecture
  • Storage architecture
  • Where quality gates and schema validation live in the architecture
  • The metadata and lineage architecture
  • Monitoring and observability
  • Key learnings

 

Chapter 4: The P in Retrieval

  • Working with source system teams
  • Responsibilities in the data retrieval phase
  • Building descriptive metadata with business
  • Escalation and incident response
  • Key learnings

 

Chapter 5: The Sample Project

  • What we build
  • The open-source stack at a glance
  • Repository layout and prerequisites
  • The three source systems
  • Classification decisions
  • Collaboration agreement: The manufacturing team
  • Defining the three Data Products with ODPS
  • The pipelines in practice
  • The shared pipeline template
  • Pipeline 1: Supplier Quality batch, in full

Apply the GAP triad of Governance, Architecture, and People to the realities of inconsistent source systems, changing schemas, unclear ownership, regulatory requirements, and uneven data quality. Profile sources before connecting to them, classify information during ingestion, establish meaningful quality gates, automate metadata capture, document lineage, and manage retention and audit logging as part of delivery rather than as an afterthought.

Design a reusable data pipeline architecture that supports batch processing, APIs, and change data capture. Evaluate connector options, orchestration patterns, storage choices, quarantine processes, and schema validation. Instead of developing a different pipeline for every use case, configure a shared template that embeds consistent controls and scales without multiplying engineering effort.

Coordinate the people behind the technology. Map responsibilities with RACI, establish productive agreements with source system teams, involve domain experts in business metadata, and define how teams respond when pipelines or quality controls fail. A detailed sample project connects the concepts using realistic source systems, YAML definitions, and an open-source stack that readers can run and examine themselves.

This book replaces fragmented practices with an implementable approach to DataOps, metadata management, data lineage, quality assurance, and dependable ingestion. Build the governed retrieval capability your organization needs to put trusted information into action.

About Mario

Mario Meir-Huber helps organizations turn scattered data initiatives into governed, business-driven Data Products. A former Head of Data and ex-Microsoftie, he has built Data Products across European companies. Mario lectures at WU Vienna, teaches on LinkedIn Learning, and speaks at events like WeAreDevelopers, Data Modeling Zone, London Tech Week, Data Science Conference, and GITEX. He co-authored The Handbook of Data Science and AI and is writing The Data Products Series. You can follow his Substack at https://futuregrade.substack.com for interesting insights on Data Products and AI.

Bestsellers

Faculty may request complimentary digital desk copies

Please complete all fields.