Math Blogs
What is RMSNorm? Root Mean Square Layer Normalization Explained
RMSNorm is a simpler and faster version of LayerNorm that keeps only the re-scaling step. It powers most modern LLMs like Llama, Mistral, and DeepSeek.
What is RoPE (Rotary Position Embedding)? The Math Behind It
RoPE (Rotary Position Embedding) gives position information to a Transformer by rotating the Q and K vectors by an angle that depends on the token's position.
What is Cross-Entropy Loss?
Cross-Entropy Loss is a number that tells us how wrong our predicted probabilities are, compared to the true answer. We will learn its math with an example.
How Does Gradient Descent Work?
Gradient Descent is an optimization algorithm that reduces a model's loss by moving its weights step by step in the direction of the steepest downward slope.
How Does Backpropagation Work? The Math Explained Step by Step
Backpropagation calculates how much each weight in a neural network contributed to the error, by sending the error backward, so that we can adjust the weights.
Why Do We Scale Attention by √dₖ? The Math Behind the Scaling Factor
We scale attention by √dₖ because dot products grow bigger as dₖ grows, which pushes softmax to extreme values. Dividing by √dₖ keeps them in a healthy range.
How Does Attention Work? The Math Behind Q, K, and V Step by Step
Attention is computed as softmax(Q x K^T / sqrt(d_k)) x V. We will learn the math behind Query, Key, and Value with a step-by-step numeric example.