Study note

Triton Tutorials

Properties

Type
Blogs
Status
待读
Domain
AI / ML
Category
GPU、CUDA 与内核优化
Source
triton-lang.org
Vault note
library/articles/ai_ml/Triton-Tutorials-b12d499789298c52.md

Summary

Triton 官方教程集,按难度提供 Vector Addition、Fused Softmax、Matrix Multiplication、Layer Normalization、Fused Attention、Group GEMM、Persistent Matmul 和 Block-scaled Matmul 等完整示例。

Highlights

把 block/tile、offset、mask、program reordering、autotuning 和持久化调度直接映射到现代 ML kernel。按 Vector Add -> Softmax -> Matmul -> Fused Attention -> Persistent Matmul 学习,可以逐步连接 CUDA 基础、GEMM、FlashAttention 与小 shape 性能问题。

Comments

Loading comments...