🤖 DeepSeek-R1: Reasoning Via Reinforcement Learning Kabir's Tech Dives podcast

Artwork

Entrepreneur Business Kabir Startup Founders Tech Podcasting Education Investors Angels

المحتوى المقدم من Kabir. يتم تحميل جميع محتويات البودكاست بما في ذلك الحلقات والرسومات وأوصاف البودكاست وتقديمها مباشرة بواسطة Kabir أو شريك منصة البودكاست الخاص بهم. إذا كنت تعتقد أن شخصًا ما يستخدم عملك المحمي بحقوق الطبع والنشر دون إذنك، فيمكنك اتباع العملية الموضحة هنا https://ar.player.fm/legal.

Kabir's Tech Dives « »
🤖 DeepSeek-R1: Reasoning via Reinforcement Learning

23d ago 13:22

مشاركة

MP3•منزل الحلقة

المحتوى المقدم من Kabir. يتم تحميل جميع محتويات البودكاست بما في ذلك الحلقات والرسومات وأوصاف البودكاست وتقديمها مباشرة بواسطة Kabir أو شريك منصة البودكاست الخاص بهم. إذا كنت تعتقد أن شخصًا ما يستخدم عملك المحمي بحقوق الطبع والنشر دون إذنك، فيمكنك اتباع العملية الموضحة هنا https://ar.player.fm/legal.

This episode details the development of DeepSeek-R1, a large language model enhanced for reasoning capabilities through reinforcement learning (RL). Two versions are described: DeepSeek-R1-Zero, trained solely with RL, and DeepSeek-R1, which incorporates a multi-stage training process including cold-start data and supervised fine-tuning to improve readability and performance. DeepSeek-R1 achieves results comparable to OpenAI's o1-1217 model on various reasoning benchmarks. Furthermore, the research explores distilling DeepSeek-R1's reasoning abilities into smaller, more efficient models, achieving strong performance despite the absence of RL in the smaller models. The authors open-source their models and findings to benefit the research community.

Support the show

Podcast:
https://kabir.buzzsprout.com
YouTube:
https://www.youtube.com/@kabirtechdives
Please subscribe and share.

… continue reading

191 حلقات

#Entrepreneur #Business #Kabir #Startup #Founders #Tech #Podcasting Education #Investors #Angels

Artwork

🤖 DeepSeek-R1: Reasoning via Reinforcement Learning

Kabir's Tech Dives

published 23d ago

مشاركة

MP3•منزل الحلقة

المحتوى المقدم من Kabir. يتم تحميل جميع محتويات البودكاست بما في ذلك الحلقات والرسومات وأوصاف البودكاست وتقديمها مباشرة بواسطة Kabir أو شريك منصة البودكاست الخاص بهم. إذا كنت تعتقد أن شخصًا ما يستخدم عملك المحمي بحقوق الطبع والنشر دون إذنك، فيمكنك اتباع العملية الموضحة هنا https://ar.player.fm/legal.

This episode details the development of DeepSeek-R1, a large language model enhanced for reasoning capabilities through reinforcement learning (RL). Two versions are described: DeepSeek-R1-Zero, trained solely with RL, and DeepSeek-R1, which incorporates a multi-stage training process including cold-start data and supervised fine-tuning to improve readability and performance. DeepSeek-R1 achieves results comparable to OpenAI's o1-1217 model on various reasoning benchmarks. Furthermore, the research explores distilling DeepSeek-R1's reasoning abilities into smaller, more efficient models, achieving strong performance despite the absence of RL in the smaller models. The authors open-source their models and findings to benefit the research community.

Support the show

Podcast:
https://kabir.buzzsprout.com
YouTube:
https://www.youtube.com/@kabirtechdives
Please subscribe and share.

… continue reading

191 حلقات

#Entrepreneur #Business #Kabir #Startup #Founders #Tech #Podcasting Education #Investors #Angels

كل الحلقات

×

مرحبًا بك في مشغل أف ام!

يقوم برنامج مشغل أف أم بمسح الويب للحصول على بودكاست عالية الجودة لتستمتع بها الآن. إنه أفضل تطبيق بودكاست ويعمل على أجهزة اندرويد والأيفون والويب. قم بالتسجيل لمزامنة الاشتراكات عبر الأجهزة.

الاستماع إلى +500 موضوع

دليل مرجعي سريع

أعلى المدونة الصوتية

SciDose بودكاست

Quizeculo كويزيكيلو

Faysalosophy Podcast | فيصلُوسُفِي بودكاست

Alkshkool بودكاست الكشكول

المحور الثاني

بودكاست كلام

Arabic News - NHK WORLD RADIO JAPAN

KBS WORLD Radio نشرة الأخبار

بزنس بالعربي (Business بالعربى )

Science Quickly

بودكاست شرفة

بداية الحكاية

Damiri | داميري

mishbilshibshib | مش بالشبشب

استمع إلى هذا العرض أثناء الاستكشاف