LLM Blogs
The Future of Software Development
The Future of Software Development is a world where AI models become so smart, so fast, and so cheap that we simply tell the computer what we want, and AI agents build, test, and ship it for us. The work of software engineers moves from writing every line of code to building, checking, and improving the systems around these AI models.
Evolution of LLM Architecture
LLM architecture evolved from RNNs that read one word at a time, to Attention, to the Transformer, to scaling, and to Mixture of Experts (MoE) that we use today.
How do Top-k and Top-p Sampling work?
Top-k Sampling keeps a fixed number of the most likely tokens, and Top-p Sampling keeps the top tokens whose total probability reaches p. The LLM picks from them.
Jev and System One Models Explained
Jev is a System One Model, a new kind of AI model that does not write text. It only makes fast decisions, with a confidence score, that our software can use directly.
Design a Real-Time Voice AI Agent
A Real-Time Voice AI Agent listens to us, understands, thinks, takes actions, and talks back in a natural voice within a fraction of a second. Let's design one.
How does Temperature control LLM output?
Temperature is a single number that controls how an LLM picks the next token. Low Temperature gives safe, predictable answers, and high gives creative ones.
What is N-gram Speculation in LLMs and How Does It Speed Up Generation?
N-gram Speculation makes an LLM write faster by guessing the next few tokens from matching text it has already seen, and then verifying them in one run.
What is Prefill-Decode Disaggregation in LLM Inference?
Prefill-Decode Disaggregation runs an LLM so that reading the prompt (prefill) and writing the answer (decode) happen on separate machines, each tuned for its job.
What is KV Cache Compression?
KV Cache Compression is a set of techniques that shrink the memory an LLM uses to remember the conversation, using quantization, token eviction, and more.
How do Attention Sinks work?
Attention sinks are the first few tokens that get a huge share of attention in an LLM. Keeping them lets StreamingLLM handle very long conversations without breaking.
How does Sliding Window Attention work?
In Sliding Window Attention, each word looks only at a fixed number of nearby words instead of all the words, which makes attention fast for long text.
How does TensorRT-LLM work?
TensorRT-LLM is an open-source library from NVIDIA that makes LLMs run as fast as possible on NVIDIA GPUs by preparing the model ahead of time.
How does an LPU work?
An LPU (Language Processing Unit) is a chip built for one job, running a trained LLM and producing text as fast as possible, by keeping the model on the chip.
What is EAGLE? Feature-Level Speculative Decoding Explained
EAGLE is a speculative decoding method that makes an LLM generate text faster, without changing its output, by guessing future tokens at the feature level.
What is Medusa? Multi-Head Speculative Decoding Explained
Medusa makes an LLM generate text 2 to 3 times faster by adding extra heads that guess several future tokens at once, and then checking them together.
How do RNNs and Transformers differ?
An RNN reads a sentence one word at a time, carrying a memory, whereas a Transformer reads all the words at once and uses attention to connect them directly.
AI Is Only as Good as Our Definition of Done
AI does its best work when we can clearly check if a task is done, and it struggles when only a human can judge it. That check is our definition of done.
How does Prefix Tuning work?
Prefix Tuning keeps a large language model frozen and trains only a small set of extra numbers, called the prefix, to adapt the model to a new task cheaply.
What is Graph Engineering?
Graph Engineering is designing an AI system as a graph, where every step is a node and every path between steps is an edge, instead of one giant prompt.
What is Loop Engineering?
Loop Engineering is designing the repeating cycle that an AI agent runs, so that it keeps making real progress on a task and stops at the right moment.
What is the Lost in the Middle Problem in LLMs and How to Fix It?
The Lost in the Middle problem is when an LLM uses the beginning and the end of a long input well, but pays very less attention to what is in the middle.
How Does LLM Watermarking Work?
LLM watermarking hides a signal inside the text a model writes. A secret key slightly nudges the word choices, and a detector with the same key finds it later.
How do LLM guardrails work?
LLM guardrails are safety checks that sit around an LLM to control what goes in and what comes out, and they stop dangerous or unwanted answers.
Encoder vs Decoder in Transformers
In Transformers, the Encoder reads the whole input in both directions to understand it, whereas the Decoder generates the output one token at a time.
What is Generative AI?
Generative AI is a type of AI that can create new things, like text, images, audio, video, and code. We will learn how it learns and creates something new.
What is Prompt Injection in LLMs and How Do We Defend Against It?
Prompt Injection is an attack where someone slips their own instructions into the text sent to an LLM, so that it follows the attacker instead of the developer.
What are Agent Skills?
An Agent Skill is a folder of instructions, and optionally scripts, that an AI agent loads by itself only when the task needs it. Let's see how Skills work.
What is MCP (Model Context Protocol)?
MCP (Model Context Protocol) is an open standard that defines one common way for AI applications to connect to outside tools and data. It was created by Anthropic.
How does fine-tuning work?
Fine-tuning is taking a model that is already trained and training it a little more on our own data, so that it becomes good at our specific task.
What is OKF (Open Knowledge Format)?
OKF (Open Knowledge Format) is an open standard for writing what an organization knows about its data as plain markdown files that any AI agent can read.
How does Cursor work?
Cursor is an AI code editor. It indexes our codebase, understands it by meaning, and helps us with Tab autocomplete, Chat, and Agent mode to write code faster.
How does context compaction work?
Context compaction shrinks the old part of a long conversation into a short summary, so that the important facts stay and the context window gets free space again.
How does Claude Code work?
Claude Code is a coding agent from Anthropic that runs in the terminal. It completes tasks by reading our code, editing files, running commands, and checking its work.
How does llama.cpp run LLMs on everyday hardware?
llama.cpp runs LLMs on everyday hardware by shrinking the model with quantization, loading it fast with memory mapping, and sharing work between CPU and GPU.
How does Chain-of-Thought (CoT) Prompting work?
Chain-of-Thought (CoT) Prompting is a technique where we ask the model to write out its reasoning steps before the final answer, which makes it more accurate.
How does Prompt Chaining work?
Prompt Chaining breaks one big task into smaller prompts, where the output of one prompt becomes the input of the next. It solves bigger tasks reliably.
Prefill vs Decode: LLM Inference Optimization
Prefill reads the whole prompt in one pass and produces the first token, whereas Decode generates the rest of the tokens one at a time using the KV cache.
How does LangGraph work?
LangGraph is a framework to build LLM applications where the work is organized as a graph of steps, with nodes, edges, and a shared state.
How does LangChain work?
LangChain is a framework that helps us build LLM applications by connecting the LLM with our data, tools, and logic through chains, memory, and agents.
How does SGLang work?
SGLang is a high-performance framework that serves LLMs to many users at the same time, as fast as possible, using ideas like RadixAttention to reuse past work.
How do Computer-Use Agents work?
A computer-use agent is an AI program that operates a computer like a human. It looks at the screen, moves the mouse, clicks, and types to finish our goal.
What is Sakana Fugu? The Technical Report Explained
Sakana Fugu is a family of AI models that work like a conductor. They decide which other AI models should work on our question and combine their answers.
How do Diffusion Language Models (DLMs) work?
A Diffusion Language Model writes text by starting from pure gibberish and cleaning it up again and again, instead of writing one word at a time from left to right.
How does vLLM work?
vLLM is a high-throughput engine for serving LLMs. It uses PagedAttention to manage the KV cache memory efficiently, so it can serve many more users at once.
LLM Inference Optimization
LLM Inference Optimization is the set of techniques, like KV Cache, Flash Attention, and Continuous Batching, that make LLMs fast and scalable in production.
How does Function Calling work in LLMs?
Function Calling lets an LLM use external tools and APIs. The model does not run the function itself. It tells our code which function to call and with what inputs.
How does GGUF work?
GGUF is a single file format that stores everything needed to run an LLM locally, like the weights, the tokenizer, and the settings, all in one file.
How does Token Streaming work?
Token Streaming sends the LLM's reply piece by piece as each piece is produced, using SSE over one open connection, instead of waiting for the whole reply.
How does Prompt Caching work?
Prompt Caching saves the work an LLM already did for a repeated part of a prompt, so that it can reuse that work next time. It makes responses faster and cheaper.
What is Continual Learning in LLMs? Solving Catastrophic Forgetting
Continual Learning is the ability of an LLM to keep learning new information over time, without forgetting what it already learned. We will see how it works.
What is Multi-Head Attention in Transformers and How Does It Work?
Multi-Head Attention runs many Self Attention operations in parallel, each focusing on a different aspect of the sentence, and combines their outputs.
What is Cross Attention in Transformers and How Does It Work?
Cross Attention is a mechanism where one sequence looks at a different sequence, using its own Queries against the Keys and Values of the other sequence.
What is Self Attention in Transformers and How Does It Work?
Self Attention lets every word in a sentence look at every other word in the same sentence to understand its meaning. It is the heart of models like BERT and GPT.
What is an AI Agent Loop?
The AI Agent Loop is the code that runs an AI Agent. It asks the LLM what to do, runs that action, feeds back the result, and repeats until the task is done.
What is AI Agent Observability? Traces, Spans, and Metrics Explained
AI Agent Observability means recording every step an AI Agent takes, like each thought and tool call, so that we can see why it behaved the way it did.
How AI Agents Communicate
AI agents communicate by sending messages to each other to share information, ask for help, and work together. We will learn the main ways and protocols.
What are AI SubAgents? Why We Need Them and How They Work
An AI SubAgent is a smaller, specialized agent that works under a main agent to handle one part of a big task, like a team member under a manager.
What is LLM Evaluation? Metrics, Benchmarks, and Methods Explained
LLM Evaluation is the process of measuring how well a Large Language Model performs on the tasks we expect it to do, using metrics, benchmarks, and more.
How to Evaluate AI Agents? Metrics, Methods, and Best Practices
AI Agent Evaluation is the process of measuring how well an AI Agent performs a task, by checking the final result, the steps it took, and the tools it used.
What is AI Orchestration? How It Works and the Common Patterns
AI Orchestration is the process of coordinating many AI parts, like LLMs, tools, data sources, and agents, so that they work together to finish a complex task.
What is LLM as a Judge? How to Use an LLM to Evaluate LLM Outputs
LLM as a Judge is a technique where we use one LLM to evaluate the output of another LLM against our criteria, and give a score or a verdict with a reason.
What are Recursive Language Models (RLMs) and How Do They Work?
A Recursive Language Model (RLM) handles very large inputs by letting the model call another language model, or itself, on smaller parts of the input.
What is Group Relative Policy Optimization (GRPO) and How Does It Work?
GRPO is a reinforcement learning algorithm to train LLMs. It generates a group of answers for the same question and teaches the model to prefer the better ones.
What is Direct Preference Optimization (DPO) and How Does It Work?
Direct Preference Optimization (DPO) trains an LLM directly on human preference data, without a separate reward model and without reinforcement learning.
What is Proximal Policy Optimization (PPO) and How Does It Work?
Proximal Policy Optimization (PPO) is a reinforcement learning algorithm that improves the policy in small, safe steps. It is used to train LLMs with RLHF.
Batch Normalization vs Layer Normalization
Batch Normalization normalizes each feature across all examples in a batch, while Layer Normalization normalizes all features within each single example.
What is RLHF? Reinforcement Learning from Human Feedback Explained
RLHF (Reinforcement Learning from Human Feedback) teaches an LLM to give the responses that humans prefer, by turning human preferences into a reward signal.
What are Large Reasoning Models (LRMs) and How Are They Different from LLMs?
A Large Reasoning Model (LRM) is an LLM that thinks step by step before it answers, producing a long chain of reasoning first. This makes it better at hard problems.
What is Continuous Batching in LLMs and How Does It Work?
Continuous Batching lets LLM servers handle many more users by replacing each finished request with a new one right away, so that the GPU is never idle.
What are Small Language Models (SLMs) and When Should We Use Them?
Small Language Models (SLMs) are language models with typically less than 10 billion parameters. They are cheaper and faster, and can run on our own devices.
What is LLM Routing? How to Send Each Query to the Right LLM
LLM Routing is the practice of choosing the right LLM for each user query, based on cost, latency, and quality, instead of sending every query to the same LLM.
What is Context Engineering?
Context Engineering is the practice of deciding what information goes into an LLM's context window, and how, so that the model can do its task reliably.
What is a Reflection Agent? How It Generates, Critiques, and Revises
A Reflection Agent is an AI Agent that writes a draft, critiques its own draft, and then writes a better version using that critique, again and again.
What is Speculative Decoding and How Does It Make LLMs Faster?
Speculative Decoding makes LLMs 2x to 3x faster. A small draft model guesses the next few tokens, and the big model verifies them all in one run, with no quality loss.
What is GraphRAG? How Knowledge Graphs Improve RAG
GraphRAG is RAG that uses a knowledge graph along with vector search to find better, more connected information before the LLM answers a question.
What is a Plan-and-Execute Agent and How Does It Work?
A Plan-and-Execute Agent is an AI Agent that first writes the full plan up front and then runs the plan step by step. Let's see how it differs from ReAct.
What is Agentic RAG? How It Works and When to Use It
Agentic RAG is a RAG system where an AI Agent decides when to search, what to search, and when to stop, instead of doing one fixed search before answering.
What is a ReAct Agent? How It Thinks and Acts, Explained
A ReAct Agent is an AI Agent built using the ReAct (Reasoning + Acting) pattern. It runs in a loop where the LLM thinks, takes an action, and observes the result.
What are Multi-Agent Systems and When Should We Use Them?
A Multi-Agent System is a setup where two or more LLM-driven agents, each with its own prompt and tools, work together on a shared task. Let's see when to use one.
How Does AI Agent Memory Work? The Memory Stack Explained
AI Agent Memory is the system that lets an LLM remember things across messages, sessions, and users. We will learn the memory stack and how it works.
What is an AI Agent? How It Works
An AI Agent is a system where an LLM takes a goal, decides the steps, uses tools, and keeps working in a loop until the goal is done. We will learn how it works.
What is RMSNorm? Root Mean Square Layer Normalization Explained
RMSNorm is a simpler and faster version of LayerNorm that keeps only the re-scaling step. It powers most modern LLMs like Llama, Mistral, and DeepSeek.
What is DeepSeek-V4 and How Does It Work? Architecture Explained
DeepSeek-V4 is a family of open Mixture-of-Experts models that supports a one-million-token context at about a tenth of the cost of DeepSeek-V3.2.
What is LoRA (Low-Rank Adaptation) and How Does It Fine-Tune LLMs?
LoRA (Low-Rank Adaptation) fine-tunes a large model without updating all its weights. It keeps the model frozen and learns a tiny pair of extra matrices.
What is RoPE (Rotary Position Embedding)? The Math Behind It
RoPE (Rotary Position Embedding) gives position information to a Transformer by rotating the Q and K vectors by an angle that depends on the token's position.
What is Grouped Query Attention (GQA) and Why Do LLMs Use It?
Grouped-Query Attention (GQA) divides the attention heads into groups, where all the heads in a group share the same Key and Value, but each has its own Query.
What is Cross-Entropy Loss?
Cross-Entropy Loss is a number that tells us how wrong our predicted probabilities are, compared to the true answer. We will learn its math with an example.
How Does Gradient Descent Work?
Gradient Descent is an optimization algorithm that reduces a model's loss by moving its weights step by step in the direction of the steepest downward slope.
What is the Feed-Forward Network in LLMs and What Does It Do?
The Feed-Forward Network is the part of every Transformer layer where each token is processed on its own, after attention. It holds most of the model's parameters.
What is Flash Attention and Why Is It So Fast?
Flash Attention computes the same attention that Transformers use, but in a much faster and more memory-efficient way on the GPU. The math and result stay the same.
What is Mixture of Experts (MoE) and How Does It Work?
Mixture of Experts (MoE) uses many small expert networks and a router that picks only a few of them for each input. This makes large models faster and cheaper.
How Does the Transformer Architecture Work?
A Transformer is a tokens-in, tokens-out machine that uses attention to understand how every word relates to every other word. It powers every modern LLM.
How Does Backpropagation Work? The Math Explained Step by Step
Backpropagation calculates how much each weight in a neural network contributed to the error, by sending the error backward, so that we can adjust the weights.
Why Do We Scale Attention by √dₖ? The Math Behind the Scaling Factor
We scale attention by √dₖ because dot products grow bigger as dₖ grows, which pushes softmax to extreme values. Dividing by √dₖ keeps them in a healthy range.
How Does Attention Work? The Math Behind Q, K, and V Step by Step
Attention is computed as softmax(Q x K^T / sqrt(d_k)) x V. We will learn the math behind Query, Key, and Value with a step-by-step numeric example.
What is Harness Engineering?
Harness Engineering is building the layer of code around an AI model that manages its inputs, outputs, tools, memory, errors, and evaluation for production.
What is Byte Pair Encoding (BPE) in LLMs?
BPE (Byte Pair Encoding) is the tokenization algorithm used by most LLMs. It keeps merging the most frequent pair of characters to build common subwords.
What is Paged Attention in LLMs and How Does It Work?
Paged Attention breaks the KV Cache memory into small fixed-size blocks called pages, which removes memory waste and lets LLMs serve many more users at once.
KV Cache in LLMs
KV Cache is a memory where an LLM saves the Key and Value of the tokens it has already processed, so that it does not compute them again for every new token.
What is Causal Masking in Attention and Why Do LLMs Need It?
Causal Masking in Attention blocks future tokens so that each token can look only at itself and the tokens before it. LLMs need it to predict the next token.