Week 8 — Lesson 4: Attention and Transformers, the intuition

3 min

Both models of this course, llama3.2 and nomic-embed-text, are Transformers. Their key step is attention: every token looks at every other token and takes a weighted mix of them. This is how a model links it to the pump ten words earlier in a NorthPeak note. This lesson gives the intuition without math, compares the two models, and explains why long texts cost so much.

Attention, heads and layers

Week 1 gave the name: the Transformer is the architecture behind llama3.2 and nomic-embed-text. Its key step is attention. Here is the intuition, without math.

Preview — the rest of the lesson is for enrolled readers.

Already enrolled with a code?

Your access is tied to your account, not to this link. Sign in with the same email you used in class: your course is waiting, no need to enter the code again.

Sign inNo account yet? Create one
This lesson is part of the “Week 8 — NLP and text representations” module

The first modules of the course are open to everyone. For the rest you have three options: buy this course once and for all, subscribe, or enter the code handed out in class.

Are you a student on this course?

The code is tied to your account: sign in or create an account and it will be applied automatically when you come back.

No account yet? Create one