Machine Learning with Python
68.6K subscribers
1.56K photos
136 videos
200 files
1.31K links
Learn Machine Learning with hands-on Python tutorials, real-world code examples, and clear explanations for researchers and developers.

Admin: @HusseinSheikho || @Hussein_Sheikho
Download Telegram
"Dive into Deep Learning" 📘🤖 is an open-source book that forms the mathematical foundation for large language models. 🧠📐

It covers linear algebra, mathematical analysis, probability theory, optimization methods, backpropagation, attention mechanisms, and transformer architectures. 🧮📉🔄

The book progressively moves from classical neural networks and convolutional neural networks to modern transformers and practical techniques used in large language models. 🚀🔗🧠

It contains over 1,000 pages 📖 and provides clear explanations, practical examples, and exercises. ✅📝 Making it one of the most comprehensive free resources for understanding the mathematical structure of modern artificial intelligence systems and language models. 🌐🔍🤖

arxiv.org/pdf/2106.11342 🔗

#DeepLearning #AI #MachineLearning #NeuralNetworks #Transformers #OpenSource

✨ Join Best TG Channels https://t.iss.one/addlist/0f6vfFbEMdAwODBk

⭐️ Join Our WhatsApp Channel https://whatsapp.com/channel/0029VaC7Weq29753hpcggW2A
❤9👍5👎1😁1
Forwarded from Data Analytics
One of the key moments when I truly understood how transformers work: 🧠✨

"Stop thinking of a transformer as a conveyor belt, where each layer transforms the output of the previous one."

"Start thinking of it as a residual flow." 🌊

Each block in a transformer has a residual connection that adds the input of the block to its output:

x' = f(x) + x


Because of this addition, the attention mechanism or MLP within the layer – the function f in the formula above – actually does NOT transform the input. Instead, it calculates the information that needs to be ADDED to the input before passing it to the next block! ➕

Imagine the main part of the transformer as a shared whiteboard. 📝 Each block reads what it needs from it and adds its own notes. All changes are additions.

Furthermore, layers can exchange information over distances. A block in the first layer can write information, and a block in the fifth layer can read it, even though there is no direct connection between them. 🔗

I owe these ideas to an older article by Anthropic about the architecture of transformers:
transformer-circuits.pub/2021/framework

#Transformers #DeepLearning #AI #MachineLearning #NeuralNetworks #Tech

✨ Join Best TG Channels https://t.iss.one/addlist/0f6vfFbEMdAwODBk

⭐️ Join Our WhatsApp Channel https://whatsapp.com/channel/0029VaC7Weq29753hpcggW2A
❤7👍1