31 subscribers
انتقل إلى وضع عدم الاتصال باستخدام تطبيق Player FM !
Training Machine Learning (ML) models on Kubernetes
Manage episode 421319868 series 3332465
In this episode of the Kubernetes Bytes podcast, Bhavin sits down with Bernie Wu, VP Strategic Partnerships and AI/CXL/Kubernetes Initiatives at Memverge. They discuss about how Kubernetes is the most popular platform to run AI model training and model inferencing jobs. The discussion dives into model training, talking about different phases of a DAG, and then talk about how Memverge can help users with efficient and cost-effective model checkpoints. The discussion goes into topics like saving costs by using spot instances, hot restart of training jobs, reclaiming unused GPU resources, etc.
Check out our website at https://kubernetesbytes.com/
Episode Sponsor: Nethopper
- Learn more about KAOPS: @nethopper.io
- For a supported-demo: info@nethopper.io
- Try the free version of KAOPS now! https://mynethopper.com/auth
Cloud Native News:
- https://www.aquasec.com/blog/linguistic-lumberjack-understanding-cve-2024-4323-in-fluent-bit/
- https://kubernetes.io/blog/2024/05/20/completing-cloud-provider-migration/
- https://thenewstack.io/introducing-aks-automatic-managed-kubernetes-for-developers/
- https://www.harness.io/blog/harness-to-acquire-split
Show Links:
- https://www.linkedin.com/in/berniewu/
- https://criu.org/Main_Page
- https://memverge.com/
- https://youtu.be/tY8YOMRuqWI?si=yB3hHqLUpYPZ-KWN
- https://youtu.be/ND4seSKpJHI?si=shh0iuA9qC-dO6eb
Timestamps:
88 حلقات
Manage episode 421319868 series 3332465
In this episode of the Kubernetes Bytes podcast, Bhavin sits down with Bernie Wu, VP Strategic Partnerships and AI/CXL/Kubernetes Initiatives at Memverge. They discuss about how Kubernetes is the most popular platform to run AI model training and model inferencing jobs. The discussion dives into model training, talking about different phases of a DAG, and then talk about how Memverge can help users with efficient and cost-effective model checkpoints. The discussion goes into topics like saving costs by using spot instances, hot restart of training jobs, reclaiming unused GPU resources, etc.
Check out our website at https://kubernetesbytes.com/
Episode Sponsor: Nethopper
- Learn more about KAOPS: @nethopper.io
- For a supported-demo: info@nethopper.io
- Try the free version of KAOPS now! https://mynethopper.com/auth
Cloud Native News:
- https://www.aquasec.com/blog/linguistic-lumberjack-understanding-cve-2024-4323-in-fluent-bit/
- https://kubernetes.io/blog/2024/05/20/completing-cloud-provider-migration/
- https://thenewstack.io/introducing-aks-automatic-managed-kubernetes-for-developers/
- https://www.harness.io/blog/harness-to-acquire-split
Show Links:
- https://www.linkedin.com/in/berniewu/
- https://criu.org/Main_Page
- https://memverge.com/
- https://youtu.be/tY8YOMRuqWI?si=yB3hHqLUpYPZ-KWN
- https://youtu.be/ND4seSKpJHI?si=shh0iuA9qC-dO6eb
Timestamps:
88 حلقات
كل الحلقات
×
1 Database as a service with Percona Everest 1:02:44


1 Increasing AI adoption using Kubernetes 52:03

1 Monolith to Microservices using Kubernetes at Guidewire 1:06:28

1 Inference in Action: Scaling Al Smarter with Inferless 55:17


1 Dagger.io Deep Dive with Co-Founder Sam Alba 1:06:24

1 Running Ray on Kubernetes with KubeRay 53:06

1 Building scalable data platforms using Data on EKS 1:02:20

1 Deploy and fine-tune LLM models on Kubernetes using KAITO 44:17

1 The business case for cloud-native and Kubernetes 54:24

1 Building the AI Hyperscaler with Kubernetes 54:56

1 Shifting Minds: Exploring OpenShift's AI Landscape 1:05:07

1 Training Machine Learning (ML) models on Kubernetes 55:29

1 The evolution of service mesh technologies 1:08:00
مرحبًا بك في مشغل أف ام!
يقوم برنامج مشغل أف أم بمسح الويب للحصول على بودكاست عالية الجودة لتستمتع بها الآن. إنه أفضل تطبيق بودكاست ويعمل على أجهزة اندرويد والأيفون والويب. قم بالتسجيل لمزامنة الاشتراكات عبر الأجهزة.