Putting It Together
- Trace one slow answer through every stage of the course, and name the module that explains each stage
- Apply fixes one at a time and combined, and read the effect on time to first word and on the full answer
- Explain why the model's writing speed is unchanged by any fix in the course
- Run the same method on a bot of your own, from measurement to a short list of changes
Six weeks later, the same team looks at the same dashboard. The first word arrives in under a second, and the support queue has gone quiet. Nobody bought a bigger model.
This module walks back through the changes that got them there, and gives you a way to repeat the process on your own bot.
The whole course on one map
Here is every stage of Pip's answer, with the module that explains it. Keep this map in mind for the rest of the module.
Stacking the fixes
Pip's team did not make one change and celebrate. They made four, one at a time, and measured after each. The interactive below does the same thing with the numbers from the earlier modules. Switch on one fix at a time, then combine them, and watch which bar moves.
Each fix is tied to the module that explains it, so you can go back and read the reasoning.
Started at 4.6 s to the first word and 5.5s for the full answer. Each fix moves one bar, and none of them touches the model's writing speed.
The order matters. The queue fix came first because it was the largest bar and the cheapest change. Index tuning came second, after the recall measurement made the trade safe. Routing and prompt trimming came last, because each one needed its own evaluation before it could ship.
Notice what did not move. The last bar, the rest of the answer, is the model writing at its own speed. The course did not make the model faster. It made the wait before the model starts shorter, and that is what the customer was waiting for.
What the numbers mean
Pip started at 4.6 seconds to the first word and 5.5 seconds for the full answer. With all four fixes on, the first word arrives at about 0.7 seconds and the full answer at about 1.6 seconds. The gap between the two in the last case is the 900 milliseconds of writing, which no fix in the course touched. In a longer answer that gap grows, and the model's writing speed matters more.
The method is the thing to keep. Measure every stage. Find the largest one. Change one thing. Measure again. The fix that moves the largest bar is usually the right next step, and a fix that moves a small bar is rarely worth doing first.
Running the method on your own bot
Here is the order to work in, for a bot of your own. It is the same order as the course.
Add a clock to every stage
Queue, embed, search, re-rank, assemble, first word and the rest of the answer. One number per stage per request, logged with the request ID. Module 1 explains how.
Read the bad afternoon
Look at p95, not the average. Pull the slowest one percent of requests and read their spans. The pattern in those requests is the pattern to fix. Module 6 explains how each stage shows its signal.
Check the queue before anything else
If queue time dominates, a faster model changes nothing. Add workers, set a pool size, or set limits on the callers. Module 9 covers the retry loop that a queue can start.
Measure recall before touching the index
Keep a set of real questions with known answers. Measure recall at your current setting, then choose the cheapest setting that meets your target. Module 3 explains the trade.
Test the router and the cache on known questions
Run a labelled set after every change to routing or caching. A fast miss is still a miss. Module 7 covers the traps.
Only then consider the model
By this point you will know whether the model is the problem. Most of the time it is not, and a model switch is a distraction until the other stages are measured.
Where this course connects
This course is one part of a larger set. The links below take you to the courses and labs that go deeper on each area, and each one is worth reading next to the module it matches.
- Production RAG, the full build from ingestion to evaluation, with the ingestion and chunking module, the retrieval module, the evaluation module and the operations module.
- Semantic Search, which builds the embeddings and index in code, including the vector database module on HNSW and IVF and the hybrid search module.
- Embeddings, in Depth, on how embedding models are trained and compared, in the embeddings course.
- Context Engineering, for how to assemble and manage what a model reads, in the context engineering course.
- Giving AI a Memory, for the design of memory across sessions, in the memory course.
- AI Agents, for the agent loop and the memory and observability modules.
- Agentic Harness Patterns, for the operational patterns that keep a slow system from becoming a broken one, including the circuit breaker, the context boundary and the trace pipeline.
- Labs: a minimal RAG system in one file, which gives you something to time, and add memory to an AI agent.
If you want the observability side in writing, the blog post on observability for AI agents covers what to track and how to set up the tooling.
Where to go from here
The model gets the blame because it is the only part of the system that talks back. Everything else is quiet until somebody times it. Time it first, and you will usually find that the slow part was never the chef.
- Every stage of a RAG answer has a module in this course, and every fix maps to one stage.
- Stacking fixes one at a time shows which bar moves and why. The largest bar is usually the right next step.
- No fix in the course changes how fast the model writes. The gains come from making the model start sooner and read less.
- The method works on any bot: measure every stage, find the largest, change one thing, measure again.
- The links in this module take you to the deeper courses and labs for each area.
- Pip's team applied the queue fix and the index fix. The first word and the full answer both drop by the same 3.35 seconds. Why does the rest-of-answer span stay unchanged, and what would have to change to shorten it?
- Your bot has a 900 ms queue, a 300 ms search and a 1.1 second first word. Which stage would you change first, which module explains it, and what evidence would tell you the change is working?
- A colleague proposes switching to a faster model before any measurement. Using this course's method, describe the one measurement that would decide the question, and what result would make the switch the right call.