DeepSeek R1: Chain Of Thought, Reinforcement Learning, And Distillation Kabir's Tech Dives podcast

Artwork

Entrepreneur Business Kabir Startup Founders Tech Podcasting Education Investors Angels

المحتوى المقدم من Kabir. يتم تحميل جميع محتويات البودكاست بما في ذلك الحلقات والرسومات وأوصاف البودكاست وتقديمها مباشرة بواسطة Kabir أو شريك منصة البودكاست الخاص بهم. إذا كنت تعتقد أن شخصًا ما يستخدم عملك المحمي بحقوق الطبع والنشر دون إذنك، فيمكنك اتباع العملية الموضحة هنا https://ar.player.fm/legal.

Kabir's Tech Dives « »
DeepSeek R1: Chain of Thought, Reinforcement Learning, and Distillation

22d ago 14:02

مشاركة

MP3•منزل الحلقة

المحتوى المقدم من Kabir. يتم تحميل جميع محتويات البودكاست بما في ذلك الحلقات والرسومات وأوصاف البودكاست وتقديمها مباشرة بواسطة Kabir أو شريك منصة البودكاست الخاص بهم. إذا كنت تعتقد أن شخصًا ما يستخدم عملك المحمي بحقوق الطبع والنشر دون إذنك، فيمكنك اتباع العملية الموضحة هنا https://ar.player.fm/legal.

DeepSeek R1, a new large language model from China, is described, highlighting three key techniques: Chain of Thought prompting to improve reasoning and self-evaluation; reinforcement learning, specifically Group Relative Policy Optimization, enabling the model to learn independently and optimize its performance without needing labeled data; and model distillation, creating smaller, more accessible versions of the model while maintaining high accuracy. These techniques allow DeepSeek R1 to achieve performance comparable to, and eventually surpassing, OpenAI's models in tasks like math, coding, and scientific reasoning. The model's innovative training methods are explained, emphasizing its efficiency and potential to democratize access to advanced AI.

Support the show

Podcast:
https://kabir.buzzsprout.com
YouTube:
https://www.youtube.com/@kabirtechdives
Please subscribe and share.

… continue reading

191 حلقات

#Entrepreneur #Business #Kabir #Startup #Founders #Tech #Podcasting Education #Investors #Angels

Artwork

DeepSeek R1: Chain of Thought, Reinforcement Learning, and Distillation

Kabir's Tech Dives

published 22d ago

مشاركة

MP3•منزل الحلقة

المحتوى المقدم من Kabir. يتم تحميل جميع محتويات البودكاست بما في ذلك الحلقات والرسومات وأوصاف البودكاست وتقديمها مباشرة بواسطة Kabir أو شريك منصة البودكاست الخاص بهم. إذا كنت تعتقد أن شخصًا ما يستخدم عملك المحمي بحقوق الطبع والنشر دون إذنك، فيمكنك اتباع العملية الموضحة هنا https://ar.player.fm/legal.

DeepSeek R1, a new large language model from China, is described, highlighting three key techniques: Chain of Thought prompting to improve reasoning and self-evaluation; reinforcement learning, specifically Group Relative Policy Optimization, enabling the model to learn independently and optimize its performance without needing labeled data; and model distillation, creating smaller, more accessible versions of the model while maintaining high accuracy. These techniques allow DeepSeek R1 to achieve performance comparable to, and eventually surpassing, OpenAI's models in tasks like math, coding, and scientific reasoning. The model's innovative training methods are explained, emphasizing its efficiency and potential to democratize access to advanced AI.

Support the show

Podcast:
https://kabir.buzzsprout.com
YouTube:
https://www.youtube.com/@kabirtechdives
Please subscribe and share.

… continue reading

191 حلقات

#Entrepreneur #Business #Kabir #Startup #Founders #Tech #Podcasting Education #Investors #Angels

كل الحلقات

×

مرحبًا بك في مشغل أف ام!

يقوم برنامج مشغل أف أم بمسح الويب للحصول على بودكاست عالية الجودة لتستمتع بها الآن. إنه أفضل تطبيق بودكاست ويعمل على أجهزة اندرويد والأيفون والويب. قم بالتسجيل لمزامنة الاشتراكات عبر الأجهزة.

الاستماع إلى +500 موضوع

دليل مرجعي سريع

أعلى المدونة الصوتية

SciDose بودكاست

Quizeculo كويزيكيلو

Faysalosophy Podcast | فيصلُوسُفِي بودكاست

Alkshkool بودكاست الكشكول

المحور الثاني

بودكاست كلام

Arabic News - NHK WORLD RADIO JAPAN

KBS WORLD Radio نشرة الأخبار

بزنس بالعربي (Business بالعربى )

Science Quickly

بودكاست شرفة

بداية الحكاية

Damiri | داميري

mishbilshibshib | مش بالشبشب

استمع إلى هذا العرض أثناء الاستكشاف