Designing Data Architecture for AI Systems
A Comprehensive Guide
Failed to add items
Add to Cart failed.
Add to Wish List failed.
Remove from wishlist failed.
Adding to library failed
Follow podcast failed
Unfollow podcast failed
Get 30 days of Standard free
Buy for $12.99
-
Narrated by:
-
Virtual Voice
-
By:
-
Shane Larson
This title uses virtual voice narration
Your AI isn't failing. Your data architecture is.
Every failed AI initiative has the same autopsy, and it's almost never the model. The model was fine. The training data lived in six systems with three definitions of "customer." The RAG pipeline indexed stale documents because nobody owned freshness. The vector store that flew in the demo fell over at production scale. The compliance team found out after it shipped. Enterprises are spending fortunes on AI and discovering, one postmortem at a time, that AI is a data architecture problem wearing a machine-learning costume.
This is the comprehensive, design-level guide to building the data foundation AI workloads actually run on. It's vendor-neutral by design — lakes, warehouses, lakehouses, vector stores, and feature stores are treated as architectural capabilities to evaluate, not products to buy, because the products will change and the architecture will outlive them. And it's organized around what AI specifically demands: the three pipelines (training, inference, and RAG) that most teams fatally conflate; data quality as the number-one failure cause; vector infrastructure past the demo cliff; governance, lineage, and privacy written from inside a regulated institution; and cost, treated as the first-class dimension it is.
It closes with what architects buy books for: complete reference architectures for a RAG platform, agent memory, and the analytics-plus-AI hybrid estate — plus honest migration roadmaps from the estate you have today to the one your AI ambitions require. It's the third volume in the "Designing X" series, completing the three-book architecture set: integration, events, and data.
What you'll learn:
- Why AI breaks the data architecture you already have — and what to do about it
- The three pipelines — training, inference, RAG — and why conflating them causes production incidents
- Vector infrastructure at scale: recall/latency/cost tradeoffs, filtering, multi-tenancy, and the re-embedding problem
- Data quality, contracts, and observability as architectural properties
- Governance, lineage, and privacy for AI that survives an audit
- Cost architecture, reference designs, and phased migration roadmaps
This book is for you if:
- You're an enterprise or data architect who now owns "make AI work here"
- You're an ML/AI engineer who needs the production data architecture around the model
- You're building RAG platforms or agents on top of existing enterprise data
- You architect AI in a regulated industry and need practices that pass examination
The model is temporary. The data architecture is the asset. Build the foundation that makes every model work.