Listen "171: Machine Learning Pipelines Are Still Data Pipelines with Sandy Ryza of Dagster"
Episode Synopsis
Highlights from this week’s conversation include:The role of an orchestrator in the lifecycle of data (1:34)Relevance of orchestration in data pipelines (00:02:45)Changes around data ops and MLOps (3:37)Data Cleaning (11:42)Overview of Dagster (13:50)Assets vs Tasks in Data Pipeline (19:15)Building a Data Pipeline with Dexter (25:40)Difference between Data Asset and Materialized Dataset (28:28)Defining Lineage and Data Assets in Dagster (29:32)The boundaries of software and organizational structures (37:25)The benefits of a unified orchestration framework (39:56)Orchestration in the development phase (45:29)The emergence of analytics engineer role (51:53)Fluidity in data pipeline and infrastructure roles (52:40)The Data Stack Show is a weekly podcast powered by RudderStack, the CDP for developers. Each week we’ll talk to data engineers, analysts, and data scientists about their experience around building and maintaining data infrastructure, delivering data and data products, and driving better outcomes across their businesses with data.RudderStack helps businesses make the most out of their customer data while ensuring data privacy and security. To learn more about RudderStack visit rudderstack.com. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
More episodes of the podcast The Data Stack Show
Re-Air: Context is King: Building Intelligent AI Analytics Platforms with Paul Blankley of Zenlytic
19/11/2025
Re-Air: From PDFs to BI and Beyond: The Future of the Data Frontend with Ryan Dolley of GoodData
05/11/2025
Re-Air: Ringing Out the Old: AI's Role in Redefining Data Teams, Tools, and Business Models
15/10/2025
ZARZA We are Zarza, the prestigious firm behind major projects in information technology.