Prof Tom, I was reviewing your video on masking, I think in one of the section you mention that CUDA avoids computing those dot products. I just want to point that in vanilla CUDA implementations, the computation with all masked tokens are performed (actually it is much faster than branching) later the mask is applied before computing the softmax. Alternatively with FlashAttention Kernels, we can see more optimization in terms of how blocks are considered and it checks whether it is all masked or unmasked or partial to determine when to compute.
I thought I will point that out.
BTW, I enjoy your lectures. Please continue with your great work
unable to access attached excel sheet. please provide access.
Its saying access denied :
https://aibyhand-my.sharepoint.com/personal/tom_aibyhand_onmicrosoft_com/_layouts/15/AccessDenied.aspx?Source=https%3A%2F%2Faibyhand%2Dmy%2Esharepoint%2Ecom%2F%3Ax%3A%2Fr%2Fpersonal%2Ftom%5Faibyhand%5Fonmicrosoft%5Fcom%2FDocuments%2F2026%2F2026%20%2D%20Seminars%2F2026%2E5%2E7%20%2D%20Gemma%204%2Fgemma%2D4%2Dpreview%2Exlsx%3Fd%3Dw7d7401da503048b99f2fc03509f2af83%26csf%3D1%26web%3D1%26e%3DxTI1D9&correlation=283314a2%2D1019%2D0000%2Dafec%2D844f875bb677&Type=item&name=16aba1aa%2Da5db%2D4728%2Db773%2De1fa82826170&listItemId=321058&listItemUniqueId=7d7401da%2D5030%2D48b9%2D9f2f%2Dc03509f2af83
I'm unable to access the excel link
Hi prof, I'm unable to access the excel sheet
Unable to access the Excel file. Please provide access
always best! 🙌
Unable to access the link like other users
Prof Tom, I was reviewing your video on masking, I think in one of the section you mention that CUDA avoids computing those dot products. I just want to point that in vanilla CUDA implementations, the computation with all masked tokens are performed (actually it is much faster than branching) later the mask is applied before computing the softmax. Alternatively with FlashAttention Kernels, we can see more optimization in terms of how blocks are considered and it checks whether it is all masked or unmasked or partial to determine when to compute.
I thought I will point that out.
BTW, I enjoy your lectures. Please continue with your great work
VJ