Listen "Training Machine Learning (ML) models on Kubernetes"
Episode Synopsis
In this episode of the Kubernetes Bytes podcast, Bhavin sits down with Bernie Wu, VP Strategic Partnerships and AI/CXL/Kubernetes Initiatives at Memverge. They discuss about how Kubernetes is the most popular platform to run AI model training and model inferencing jobs. The discussion dives into model training, talking about different phases of a DAG, and then talk about how Memverge can help users with efficient and cost-effective model checkpoints. The discussion goes into topics like saving costs by using spot instances, hot restart of training jobs, reclaiming unused GPU resources, etc. Check out our website at https://kubernetesbytes.com/ Episode Sponsor: Nethopper Learn more about KAOPS: @nethopper.io For a supported-demo: [email protected] Try the free version of KAOPS now! https://mynethopper.com/authCloud Native News: https://www.aquasec.com/blog/linguistic-lumberjack-understanding-cve-2024-4323-in-fluent-bit/ https://kubernetes.io/blog/2024/05/20/completing-cloud-provider-migration/ https://thenewstack.io/introducing-aks-automatic-managed-kubernetes-for-developers/ https://www.harness.io/blog/harness-to-acquire-splitShow Links: https://www.linkedin.com/in/berniewu/ https://criu.org/Main_Page https://memverge.com/ https://youtu.be/tY8YOMRuqWI?si=yB3hHqLUpYPZ-KWN https://youtu.be/ND4seSKpJHI?si=shh0iuA9qC-dO6ebTimestamps: 01:04 Cloud Native News 08:47 Interview with Bernie 51:40 Key takeaways
More episodes of the podcast Kubernetes Bytes
Database as a service with Percona Everest
03/03/2025
KubeCon NA 2024 News Recap
18/12/2024
Increasing AI adoption using Kubernetes
06/12/2024
Container security with Wiz
07/10/2024
Dagger.io Deep Dive with Co-Founder Sam Alba
23/09/2024
Running Ray on Kubernetes with KubeRay
05/09/2024
ZARZA We are Zarza, the prestigious firm behind major projects in information technology.