Data Engineering Fundamentals for ML Pipelines Audiolibro Por Jordan O'Neal arte de portada

Data Engineering Fundamentals for ML Pipelines

A Practical Path from Spreadsheets to Production ML Data for New Engineers and Career Switchers

Muestra de Voz Virtual

Obtén 30 días de Standard gratis

Prime logotipo ¿Miembros Prime? ¿Nuevo en Audible? Recibes una prueba de 2 meses en su lugar.
$8.99 al mes después de que termine la prueba. Cancela en cualquier momento
Pruébalo por $0.00
Más opciones de compra

Data Engineering Fundamentals for ML Pipelines

De: Jordan O'Neal
Narrado por: Virtual Voice
Pruébalo por $0.00

$8.99 al mes después de 30 días. Cancela en cualquier momento.

Compra ahora por $11.99

Compra ahora por $11.99

Background images

Este título utiliza narración de voz virtual

Voz Virtual es una narración generada por computadora para audiolibros..
This book includes all the necessary code samples and diagrams to learn the ML process.
Inside this book, readers will learn how to:
  • Shift from query thinking to pipeline thinking — understanding what a human does unconsciously when running a query manually, and how to encode every one of those checks into a pipeline that runs safely without supervision
  • Design ingestion steps that land raw data correctly — using watermarks, incremental load patterns, and raw-landing discipline to prevent partial loads and late-arriving data from silently corrupting downstream models
  • Build transformation layers with explicit contracts — cleaning, typing, and shaping data with declared input-output boundaries that prevent the silent cross-team failures most ML pipeline incidents trace back to
  • Gate every pipeline stage with assertion-based validation — using row-count checks, null-rate thresholds, range assertions, and schema drift detection to stop bad data before it reaches the feature table or the model
  • Orchestrate pipelines as DAGs with safe backfill and idempotent writes — so that re-running any step twice produces the same result as running it once, and a failed job never leaves data in a half-written state
  • Engineer features with point-in-time correctness — preventing training-serving skew by computing features from only the data that would have been available at prediction time, and designing feature stores that serve both offline training and online inference from a consistent source
  • Design a monitoring configuration that a half-asleep teammate can act on — separating urgent alerts from informational ones, setting thresholds that fire when they should and stay quiet when they should, and writing runbooks that answer the three questions every on-call engineer asks at 2 a.m.: what broke, why, and what to do
  • Walk a complete daily churn feature pipeline from blank page to monitored production system — applying all six organs in sequence so the full arc from idea to working pipeline is concrete rather than theoretical
The book is organized around a six-organ framework that describes every production ML data pipeline in the same vocabulary: ingestion, transformation, validation, storage, serving, and monitoring. That framework appears in Chapter 2 and every subsequent chapter refers back to it — making the book work both as a sequential read and as a reference when you return to a specific layer later. The eleven chapters build from the pipeline mental model through each organ in sequence, closing with a fully worked end-to-end example that applies all six organs to a real daily pipeline with real design decisions.
The tools named throughout — Airflow, dbt, Great Expectations, Snowflake, BigQuery, Databricks, Delta Lake, Iceberg, Feast, Tecton — are examples of patterns, not endorsements. The underlying concepts are the spine; the named tools are one concrete instance of each pattern. Tools evolve. The six-organ framework, idempotency requirement, point-in-time discipline, and monitoring-with-runbooks habit are more durable — and they are what this book teaches.
For the SQL analyst automating their first pipeline, the early-career data engineer who wants to raise their engineering quality threshold, and the ML practitioner who has realized that most of their production problems are data problems, this is the foundation the field requires before the pipeline deserves to run unattended.
Ciencia de Datos Informática Programación
adbl_web_anon_alc_button_suppression_t1
Todavía no hay opiniones