Pranav Srivastava
Learning tracks
publishedIntermediate10 modules2h 14m

Why Your RAG Chatbot Is Slow

Find where a RAG chatbot really loses its time, then fix the right stage. Embeddings, HNSW and IVF, re-ranking, the context window, memory, queues and ingestion, followed through one customer question with interactive timers.

RAGLatencyVector SearchHNSWIVFContext WindowMemoryObservability

Course modules

01

The Four and a Half Seconds

A customer waits four and a half seconds for a chatbot to start answering. This module times every stage of that wait, explains percentiles and time to first token, and shows why the model is rarely the first suspect.

14m
02

What Search Really Does: Embeddings and Similarity

Before you can make vector search fast, you need to know what it compares. Embeddings turn meaning into coordinates, similarity measures how close two meanings are, and exact search is the honest baseline.

12m
03

Approximate Search: HNSW, IVF and Recall

Exact search is too slow at scale, so vector databases take shortcuts. This module explains HNSW and IVF from first principles, shows the trade-off through recall and latency, and gives a method for choosing a setting.

16m
04

Chunks, Re-ranking and the Context Window

What the search actually returns, why a second and slower reader improves the top results, and how the model's context window and prefill decide when the first word can appear.

12m
05

Context and Memory: What the Model Reads Each Turn

A language model remembers nothing between calls. The application decides what the model reads each turn, from instructions and passages to conversation history and customer facts. This module shows how to build that prompt, what each memory strategy costs, and how long-term memory fits in.

16m
06

The Online Path, Stage by Stage

Walk the request path from the moment a question arrives to the last streamed word. Each stage gets its own section, covering how it becomes slow in practice, the signs to look for and the first fix to try.

12m
07

Shelves, Routes and the Cheap Fixes

Split the index so a question searches one shelf, not the whole library. Then use caching, parallel calls and warm clients, and learn what each of them hides.

14m
08

The Offline Pipeline: When the Upload Is Slow

Parsing, chunking, embedding and storing run when documents arrive, not when people ask. This module follows a 400-page upload through the line, explains why the slowest stage sets the pace, and shows how to re-index only what changed.

12m
09

Who Pays, and Where to Look First

Slowness costs the person waiting, the support team, the product, the bill and the answer's quality. This module traces those costs, shows how a slow system feeds its own failure, separates the owners of each span, and gives a triage table for the first look.

14m
10

Putting It Together

Every fix in this course lands on one timeline. This module stacks them on Pip's four and a half seconds, shows what each one buys and where it is explained, and gives a method for running the same process on your own bot.

12m
All tracksQuestions? Get in touch →