Data Engineering Fundamentals for ML Pipelines
A Practical Path from Spreadsheets to Production ML Data for New Engineers and Career Switchers
No se pudo agregar al carrito
Solo puedes tener X títulos en el carrito para realizar el pago.
Add to Cart failed.
Por favor prueba de nuevo más tarde
Error al Agregar a Lista de Deseos.
Por favor prueba de nuevo más tarde
Error al eliminar de la lista de deseos.
Por favor prueba de nuevo más tarde
Error al añadir a tu biblioteca
Por favor intenta de nuevo
Error al seguir el podcast
Intenta nuevamente
Error al dejar de seguir el podcast
Intenta nuevamente
Obtén 30 días de Standard gratis
¿Miembros Prime? ¿Nuevo en Audible? Recibes una prueba de 2 meses en su lugar.
$8.99 al mes después de que termine la prueba. Cancela en cualquier momento
Compra ahora por $11.99
-
Narrado por:
-
Virtual Voice
-
De:
-
Jordan O'Neal
Este título utiliza narración de voz virtual
Voz Virtual es una narración generada por computadora para audiolibros..
Inside this book, readers will learn how to:
- Shift from query thinking to pipeline thinking — understanding what a human does unconsciously when running a query manually, and how to encode every one of those checks into a pipeline that runs safely without supervision
- Design ingestion steps that land raw data correctly — using watermarks, incremental load patterns, and raw-landing discipline to prevent partial loads and late-arriving data from silently corrupting downstream models
- Build transformation layers with explicit contracts — cleaning, typing, and shaping data with declared input-output boundaries that prevent the silent cross-team failures most ML pipeline incidents trace back to
- Gate every pipeline stage with assertion-based validation — using row-count checks, null-rate thresholds, range assertions, and schema drift detection to stop bad data before it reaches the feature table or the model
- Orchestrate pipelines as DAGs with safe backfill and idempotent writes — so that re-running any step twice produces the same result as running it once, and a failed job never leaves data in a half-written state
- Engineer features with point-in-time correctness — preventing training-serving skew by computing features from only the data that would have been available at prediction time, and designing feature stores that serve both offline training and online inference from a consistent source
- Design a monitoring configuration that a half-asleep teammate can act on — separating urgent alerts from informational ones, setting thresholds that fire when they should and stay quiet when they should, and writing runbooks that answer the three questions every on-call engineer asks at 2 a.m.: what broke, why, and what to do
- Walk a complete daily churn feature pipeline from blank page to monitored production system — applying all six organs in sequence so the full arc from idea to working pipeline is concrete rather than theoretical
The tools named throughout — Airflow, dbt, Great Expectations, Snowflake, BigQuery, Databricks, Delta Lake, Iceberg, Feast, Tecton — are examples of patterns, not endorsements. The underlying concepts are the spine; the named tools are one concrete instance of each pattern. Tools evolve. The six-organ framework, idempotency requirement, point-in-time discipline, and monitoring-with-runbooks habit are more durable — and they are what this book teaches.
For the SQL analyst automating their first pipeline, the early-career data engineer who wants to raise their engineering quality threshold, and the ML practitioner who has realized that most of their production problems are data problems, this is the foundation the field requires before the pipeline deserves to run unattended.
adbl_web_anon_alc_button_suppression_t1
Todavía no hay opiniones