System Design Blogs
How do Browser Agents work?
Browser Agents are AI agents that use a web browser like a human does. They look at a web page, decide what to click or type, do it, look at the new page, and repeat until the task is finished.
How do Query Rewriting and Multi-Query Retrieval work?
Query Rewriting and Multi-Query Retrieval are two ways of improving the search step of an AI system. Query Rewriting uses an LLM to turn the user's messy question into one clear search query, and Multi-Query Retrieval uses an LLM to turn the question into several different search queries, searches with all of them, and merges the results.
How does Tree of Thoughts work?
Tree of Thoughts is a way of making an LLM solve a hard problem by trying many possible next steps, checking which ones look good, and going deeper only on the good ones, just like exploring the branches of a tree.
How does Contextual Retrieval work?
Contextual Retrieval is a way to improve RAG by adding a short note to every chunk of a document before we store it. The note explains where the chunk comes from, so the system can find the right chunk even when the chunk alone does not say enough.
How does Ollama work?
Ollama is a free tool that lets us download and run LLMs on our own computer with one simple command. It works by running a small server in the background that downloads compressed model files, loads them into memory, and uses a fast engine to generate text when we send it a prompt.
Stop Tokens in LLMs
A Stop Token is a special token in the vocabulary of an LLM that means the text ends here. When the model produces it, the generation loop stops writing.
How do Top-k and Top-p Sampling work?
Top-k Sampling keeps a fixed number of the most likely tokens, and Top-p Sampling keeps the top tokens whose total probability reaches p. The LLM picks from them.
Design a Real-Time Voice AI Agent
A Real-Time Voice AI Agent listens to us, understands, thinks, takes actions, and talks back in a natural voice within a fraction of a second. Let's design one.
How does Temperature control LLM output?
Temperature is a single number that controls how an LLM picks the next token. Low Temperature gives safe, predictable answers, and high gives creative ones.
What is N-gram Speculation in LLMs and How Does It Speed Up Generation?
N-gram Speculation makes an LLM write faster by guessing the next few tokens from matching text it has already seen, and then verifying them in one run.
What is Prefill-Decode Disaggregation in LLM Inference?
Prefill-Decode Disaggregation runs an LLM so that reading the prompt (prefill) and writing the answer (decode) happen on separate machines, each tuned for its job.
How to Chunk Documents for RAG? Chunking Strategies Explained
Chunking for RAG means cutting a big document into smaller pieces so that the AI can find the right piece at the right time. We will learn the best strategies.
What is Vectorless RAG? RAG Without Embeddings or a Vector Database
Vectorless RAG answers questions from our documents without embeddings or a vector database. The LLM reads the document structure and decides which part to open.
How does TensorRT-LLM work?
TensorRT-LLM is an open-source library from NVIDIA that makes LLMs run as fast as possible on NVIDIA GPUs by preparing the model ahead of time.
How does an LPU work?
An LPU (Language Processing Unit) is a chip built for one job, running a trained LLM and producing text as fast as possible, by keeping the model on the chip.
AI Is Only as Good as Our Definition of Done
AI does its best work when we can clearly check if a task is done, and it struggles when only a human can judge it. That check is our definition of done.
What is Graph Engineering?
Graph Engineering is designing an AI system as a graph, where every step is a node and every path between steps is an edge, instead of one giant prompt.
What is Loop Engineering?
Loop Engineering is designing the repeating cycle that an AI agent runs, so that it keeps making real progress on a task and stops at the right moment.
What is the Lost in the Middle Problem in LLMs and How to Fix It?
The Lost in the Middle problem is when an LLM uses the beginning and the end of a long input well, but pays very less attention to what is in the middle.
How Does LLM Watermarking Work?
LLM watermarking hides a signal inside the text a model writes. A secret key slightly nudges the word choices, and a detector with the same key finds it later.
How do LLM guardrails work?
LLM guardrails are safety checks that sit around an LLM to control what goes in and what comes out, and they stop dangerous or unwanted answers.
Cloud vs On-device Model Deployment
In Cloud Deployment, the AI model runs on a server and the device asks it over the internet. In On-device Deployment, the model runs inside the device itself.
What is Prompt Injection in LLMs and How Do We Defend Against It?
Prompt Injection is an attack where someone slips their own instructions into the text sent to an LLM, so that it follows the attacker instead of the developer.
What are Embeddings?
An embedding is a list of numbers that represents the meaning of something, so that things with similar meaning get similar numbers. It powers AI search.
What are Agent Skills?
An Agent Skill is a folder of instructions, and optionally scripts, that an AI agent loads by itself only when the task needs it. Let's see how Skills work.
What is OKF (Open Knowledge Format)?
OKF (Open Knowledge Format) is an open standard for writing what an organization knows about its data as plain markdown files that any AI agent can read.
How does Semantic Search work?
Semantic Search finds results based on meaning, not exact words. It turns text into embeddings and returns the items whose meaning is closest to the query.
How does Cursor work?
Cursor is an AI code editor. It indexes our codebase, understands it by meaning, and helps us with Tab autocomplete, Chat, and Agent mode to write code faster.
How does context compaction work?
Context compaction shrinks the old part of a long conversation into a short summary, so that the important facts stay and the context window gets free space again.
How does Claude Code work?
Claude Code is a coding agent from Anthropic that runs in the terminal. It completes tasks by reading our code, editing files, running commands, and checking its work.
How does llama.cpp run LLMs on everyday hardware?
llama.cpp runs LLMs on everyday hardware by shrinking the model with quantization, loading it fast with memory mapping, and sharing work between CPU and GPU.
How does Chain-of-Thought (CoT) Prompting work?
Chain-of-Thought (CoT) Prompting is a technique where we ask the model to write out its reasoning steps before the final answer, which makes it more accurate.
How does Prompt Chaining work?
Prompt Chaining breaks one big task into smaller prompts, where the output of one prompt becomes the input of the next. It solves bigger tasks reliably.
How does Semantic Caching work?
Semantic Caching is a cache that matches questions by their meaning instead of their exact words, so an AI app can reuse past answers for similar questions.
How does Hybrid Search work?
Hybrid Search combines keyword search and semantic search, and merges their results into one final ranked list, so that we get the best of both.
How does HyDE work in RAG?
HyDE (Hypothetical Document Embeddings) is a RAG technique where the AI first writes a fake answer to the question, and then we search using that fake answer.
Prefill vs Decode: LLM Inference Optimization
Prefill reads the whole prompt in one pass and produces the first token, whereas Decode generates the rest of the tokens one at a time using the KV cache.
How does a GPU work for Deep Learning?
A GPU is a chip that does a huge number of simple calculations at the same time. Deep learning is mostly matrix multiplication, so GPUs make it very fast.
How does LangGraph work?
LangGraph is a framework to build LLM applications where the work is organized as a graph of steps, with nodes, edges, and a shared state.
How does LangChain work?
LangChain is a framework that helps us build LLM applications by connecting the LLM with our data, tools, and logic through chains, memory, and agents.
How does SGLang work?
SGLang is a high-performance framework that serves LLMs to many users at the same time, as fast as possible, using ideas like RadixAttention to reuse past work.
How does Approximate Nearest Neighbor (ANN) search work?
Approximate Nearest Neighbor (ANN) Search finds items that are very close to the most similar one, much faster, by allowing a tiny chance of not being exact.
How does a Google TPU work?
A TPU (Tensor Processing Unit) is a chip made by Google for machine learning. It uses a systolic array to do matrix multiplication very fast and efficiently.
How do Image Embeddings work?
An image embedding is a list of numbers that represents the meaning of an image. Similar images get similar numbers, so we can compare and search images easily.
How do Computer-Use Agents work?
A computer-use agent is an AI program that operates a computer like a human. It looks at the screen, moves the mouse, clicks, and types to finish our goal.
How does an Embedding Cache work?
An Embedding Cache stores the embeddings we have already computed, so that we reuse them for the same text instead of computing them again. It saves money and time.
How does vLLM work?
vLLM is a high-throughput engine for serving LLMs. It uses PagedAttention to manage the KV cache memory efficiently, so it can serve many more users at once.
LLM Inference Optimization
LLM Inference Optimization is the set of techniques, like KV Cache, Flash Attention, and Continuous Batching, that make LLMs fast and scalable in production.
How does GGUF work?
GGUF is a single file format that stores everything needed to run an LLM locally, like the weights, the tokenizer, and the settings, all in one file.
How does Token Streaming work?
Token Streaming sends the LLM's reply piece by piece as each piece is produced, using SSE over one open connection, instead of waiting for the whole reply.
How does Prompt Caching work?
Prompt Caching saves the work an LLM already did for a repeated part of a prompt, so that it can reuse that work next time. It makes responses faster and cheaper.
How does a Reranker work?
A Reranker is a model that takes a list of documents and reorders them, putting the most relevant ones at the top for a given question. It makes RAG more accurate.
How does a Vector Database work?
A Vector Database stores data as lists of numbers and helps us find items that are similar in meaning, not just items that match exactly. It powers AI search.
What is Speculative Decoding and How Does It Make LLMs Faster?
Speculative Decoding makes LLMs 2x to 3x faster. A small draft model guesses the next few tokens, and the big model verifies them all in one run, with no quality loss.
Android Push Notification Flow using FCM
In the FCM flow, the app gets a token from FCM and sends it to our backend. The backend then asks FCM to deliver the notification to that token.
Android System Design Interviews
Android System Design is the process of defining the client components, API requirements, and database tables of an app to meet the given requirements.
Write-Ahead Logging (WAL)
Write-Ahead Logging (WAL) is a database technique where every change is first written to an append-only log before the database files, to keep data safe in a crash.
Composite Index in Database
A Composite Index is an index created using multiple columns of a table. It speeds up queries on those columns, and the order of the columns matters a lot.
Database Normalization vs Denormalization
Database normalization focuses on reducing data duplication and keeping data correct, whereas denormalization focuses on making the queries faster.
Evolution of HTTP
HTTP evolved from HTTP 1.0 to 1.1, 2.0, and 3.0, where each version solved the problem of the previous one, like new TCP connections and HOL blocking.
Internals of RESP - Redis Serialization Protocol
RESP (Redis Serialization Protocol) is the text-based language that Redis clients and servers use to talk. The first byte of each message tells the data type.
HTTP Request vs HTTP Long-Polling vs WebSocket vs Server-Sent Events
HTTP Request asks once, Long-Polling waits until data is ready, WebSocket keeps a two-way connection open, and SSE lets only the server push data to the client.
What is System Design?
System Design is the process of defining the components, APIs, and database tables of a system to satisfy the given functional and non-functional requirements.
How do Voice And Video Call Work?
Voice and video calls work using WebRTC. The two clients first exchange details through signaling, then connect peer-to-peer using STUN, or via a TURN server.