Assemble a Transformer block and apply a causal mask that stops a token from seeing the future.
Proof of success: explain "Transformer, positions and the future mask", complete "The Visibility Triangle" and name the limit of the analogy.
To finish a story card by card, Imran builds a model that processes several positions together. Without position information, attention receives no clue distinguishing the first spot from the fourth; Imran therefore adds that clue. In a causal generative model, card number 3 may only consult 1, 2 and 3. A mask makes future positions inaccessible. After attention come the residual connection, normalization and a feed-forward network.
Preview — the rest of the lesson is for enrolled readers.
Your access is tied to your account, not to this link. Sign in with the same email you used in class: your course is waiting, no need to enter the code again.
Sign inNo account yet? Create oneThe first modules of the course are open to everyone. For the rest you have three options: buy this course once and for all, subscribe, or enter the code handed out in class.