Lab · WebGPU
The 207 kernels, live on your GPU
A kernel is the smallest program that runs on a GPU: a single operation on numbers, and everything else in AI is combining them. @huggingface/kernels exposes 207. Tap any of them and watch it run on your card: the curve of an activation, the matrices of a multiplication, the effect on an image.
Arithmetic and trigonometry 33
Element-wise operations: add, multiply, powers, roots, sine and cosine. The basis for blending layers, adjusting brightness or contrast and warping coordinates. Example: brightening a photo adds a constant to every pixel; a fade multiplies by a number that goes down.
Activations 28
Non-linear curves — sigmoid, tanh, softmax, gelu. In a network they decide which neuron “fires”; on an image they are tone and enhancement curves. Example: the sigmoid turns any number into a probability between 0 and 1 — the last layer of a classifier.
Convolution and pooling 18
Sliding a filter over the image: sharpening, blurring, edges, emboss. And lowering resolution with pooling. It is the heart of a convolutional network. Example: edge detection is a 3×3 convolution; halving an image while keeping what matters, a pooling.
Algebra (matmul/gemm) 12
Matrix multiplication: dense layers, projections and color transforms by matrix. Includes the quantized versions that run an LLM. Example: every dense layer of a network is a matrix multiplication; converting an image to grayscale, a 1×3 matrix.
Normalization 12
Rescaling activations to a stable range (zero mean, unit variance) so a deep network neither explodes nor dies out. Example: every block of a transformer goes through LayerNorm — a GPT-style model does it hundreds of times per token.
Attention and transformer 13
The building blocks of a language model: attention, cached attention, rotary embeddings, mixture of experts. With these you can assemble LLM inference by hand. Example: attention decides, for each word, how much it looks at the others — the core mechanism of a GPT-style model.
Quantization and types 7
Compressing weights to int4/int8 and converting types. MatMulNBits runs a quantized model without decompressing it in memory. Example: MatMulNBits runs a 7B model in 4 bits that would not fit on the card in 16 bits.
Reduction and statistics 16
Collapsing an axis: mean, max, sum, top-k. Used for global pooling, classification and histograms. Example: ArgMax over the last layer returns the predicted class; ReduceMean, the average of an activation map.
Logic and comparison 20
Masks and conditional selection. With Greater and Where you get chroma key, thresholding and cutouts by color. Example: removing a green screen compares every pixel and picks with Where between the photo and the new background.
Shape and tensors 32
Rearranging without computing: cropping, transposing, concatenating, packing, pixel-shuffle. The plumbing that connects one operation to the next. Example: turning an image from [height, width, color] into [color, height, width] for the network is a Transpose.
Resizing and sampling 4
Scaling and warping the pixel grid. Example: video upscalers use Resize; Stable Diffusion uses GridSample to warp attention maps over the image.
Signal (DFT/STFT) 5
Fourier transforms and windows. Spectrograms and audio processing on the GPU. Example: the DFT turns an audio fragment into its frequencies — the first step of a speech recognizer.
Sequence and control 6
Recurrent cells (RNN, LSTM, GRU) and control flow (Loop, Scan, If). Models that advance step by step. Example: an LSTM processes a time series while remembering what came before; Loop repeats a block N times.
Other 1
Standalone operators that do not fit a clear family.