Math Blogs

What is RMSNorm? Root Mean Square Layer Normalization Explained

What is RMSNorm? Root Mean Square Layer Normalization Explained

RMSNorm is a simpler and faster version of LayerNorm that keeps only the re-scaling step. It powers most modern LLMs like Llama, Mistral, and DeepSeek.

What is RoPE (Rotary Position Embedding)? The Math Behind It

What is RoPE (Rotary Position Embedding)? The Math Behind It

RoPE (Rotary Position Embedding) gives position information to a Transformer by rotating the Q and K vectors by an angle that depends on the token's position.

What is Cross-Entropy Loss?

What is Cross-Entropy Loss?

Cross-Entropy Loss is a number that tells us how wrong our predicted probabilities are, compared to the true answer. We will learn its math with an example.

How Does Gradient Descent Work?

How Does Gradient Descent Work?

Gradient Descent is an optimization algorithm that reduces a model's loss by moving its weights step by step in the direction of the steepest downward slope.

How Does Backpropagation Work? The Math Explained Step by Step

How Does Backpropagation Work? The Math Explained Step by Step

Backpropagation calculates how much each weight in a neural network contributed to the error, by sending the error backward, so that we can adjust the weights.

Why Do We Scale Attention by √dₖ? The Math Behind the Scaling Factor

Why Do We Scale Attention by √dₖ? The Math Behind the Scaling Factor

We scale attention by √dₖ because dot products grow bigger as dₖ grows, which pushes softmax to extreme values. Dividing by √dₖ keeps them in a healthy range.

How Does Attention Work? The Math Behind Q, K, and V Step by Step

How Does Attention Work? The Math Behind Q, K, and V Step by Step

Attention is computed as softmax(Q x K^T / sqrt(d_k)) x V. We will learn the math behind Query, Key, and Value with a step-by-step numeric example.