Why Your RAG Chatbot Is Slow
Find where a RAG chatbot really loses its time, then fix the right stage. Embeddings, HNSW and IVF, re-ranking, the context window, memory, queues and ingestion, followed through one customer question with interactive timers.
Course modules
01The Four and a Half Seconds
A customer waits four and a half seconds for a chatbot to start answering. This module times every stage of that wait, explains percentiles and time to first token, and shows why the model is rarely the first suspect.
What Search Really Does: Embeddings and Similarity
Before you can make vector search fast, you need to know what it compares. Embeddings turn meaning into coordinates, similarity measures how close two meanings are, and exact search is the honest baseline.
Approximate Search: HNSW, IVF and Recall
Exact search is too slow at scale, so vector databases take shortcuts. This module explains HNSW and IVF from first principles, shows the trade-off through recall and latency, and gives a method for choosing a setting.
Chunks, Re-ranking and the Context Window
What the search actually returns, why a second and slower reader improves the top results, and how the model's context window and prefill decide when the first word can appear.
Context and Memory: What the Model Reads Each Turn
A language model remembers nothing between calls. The application decides what the model reads each turn, from instructions and passages to conversation history and customer facts. This module shows how to build that prompt, what each memory strategy costs, and how long-term memory fits in.
The Online Path, Stage by Stage
Walk the request path from the moment a question arrives to the last streamed word. Each stage gets its own section, covering how it becomes slow in practice, the signs to look for and the first fix to try.
Shelves, Routes and the Cheap Fixes
Split the index so a question searches one shelf, not the whole library. Then use caching, parallel calls and warm clients, and learn what each of them hides.
The Offline Pipeline: When the Upload Is Slow
Parsing, chunking, embedding and storing run when documents arrive, not when people ask. This module follows a 400-page upload through the line, explains why the slowest stage sets the pace, and shows how to re-index only what changed.
Who Pays, and Where to Look First
Slowness costs the person waiting, the support team, the product, the bill and the answer's quality. This module traces those costs, shows how a slow system feeds its own failure, separates the owners of each span, and gives a triage table for the first look.
Putting It Together
Every fix in this course lands on one timeline. This module stacks them on Pip's four and a half seconds, shows what each one buys and where it is explained, and gives a method for running the same process on your own bot.
Where to go next
Agentic AI Harness Patterns
What is an AI agent harness, and why do ten named patterns keep it out of trouble? A field guide built on real incidents, a running example, and an interactive playground for every pattern.
AI Ops
The difference between a demo and a product that survives real users: monitoring, cost control, security, guardrails, and rollback. The operational layer that keeps AI systems safe, affordable, and trustworthy in production.