Pranav Srivastava
Why Your RAG Chatbot Is Slow/Putting It Together10/10

10 lessons

0/10 done
Lesson 10 of 10·12 min·Intermediate
1012 min

Putting It Together

By the end of this chapter, you'll be able to:
  • Trace one slow answer through every stage of the course, and name the module that explains each stage
  • Apply fixes one at a time and combined, and read the effect on time to first word and on the full answer
  • Explain why the model's writing speed is unchanged by any fix in the course
  • Run the same method on a bot of your own, from measurement to a short list of changes

Six weeks later, the same team looks at the same dashboard. The first word arrives in under a second, and the support queue has gone quiet. Nobody bought a bigger model.

This module walks back through the changes that got them there, and gives you a way to repeat the process on your own bot.

The whole course on one map

Here is every stage of Pip's answer, with the module that explains it. Keep this map in mind for the rest of the module.

The course, stage by stage
Each module owns one part of Pip's four and a half seconds

Stacking the fixes

Pip's team did not make one change and celebrate. They made four, one at a time, and measured after each. The interactive below does the same thing with the numbers from the earlier modules. Switch on one fix at a time, then combine them, and watch which bar moves.

Each fix is tied to the module that explains it, so you can go back and read the reasoning.

Pip's stack of fixes. Switch them on one at a time, then all together.
Waiting in line
2600 ms
Embed the query
25 ms
Vector search
1400 ms
Re-rank
180 ms
First word
350 ms
Rest of the answer
900 ms
First word
4.55 s
Full answer
5.46 s

Started at 4.6 s to the first word and 5.5s for the full answer. Each fix moves one bar, and none of them touches the model's writing speed.

The order matters. The queue fix came first because it was the largest bar and the cheapest change. Index tuning came second, after the recall measurement made the trade safe. Routing and prompt trimming came last, because each one needed its own evaluation before it could ship.

Notice what did not move. The last bar, the rest of the answer, is the model writing at its own speed. The course did not make the model faster. It made the wait before the model starts shorter, and that is what the customer was waiting for.

What the numbers mean

Pip started at 4.6 seconds to the first word and 5.5 seconds for the full answer. With all four fixes on, the first word arrives at about 0.7 seconds and the full answer at about 1.6 seconds. The gap between the two in the last case is the 900 milliseconds of writing, which no fix in the course touched. In a longer answer that gap grows, and the model's writing speed matters more.

The method is the thing to keep. Measure every stage. Find the largest one. Change one thing. Measure again. The fix that moves the largest bar is usually the right next step, and a fix that moves a small bar is rarely worth doing first.

Running the method on your own bot

Here is the order to work in, for a bot of your own. It is the same order as the course.

Add a clock to every stage

Queue, embed, search, re-rank, assemble, first word and the rest of the answer. One number per stage per request, logged with the request ID. Module 1 explains how.

Read the bad afternoon

Look at p95, not the average. Pull the slowest one percent of requests and read their spans. The pattern in those requests is the pattern to fix. Module 6 explains how each stage shows its signal.

Check the queue before anything else

If queue time dominates, a faster model changes nothing. Add workers, set a pool size, or set limits on the callers. Module 9 covers the retry loop that a queue can start.

Measure recall before touching the index

Keep a set of real questions with known answers. Measure recall at your current setting, then choose the cheapest setting that meets your target. Module 3 explains the trade.

Print the prompt size next to the first word

If the prompt is big, cut it to the best few passages and measure again. Check whether the instructions can be cached. Modules 4 and 5 explain the budget.

Test the router and the cache on known questions

Run a labelled set after every change to routing or caching. A fast miss is still a miss. Module 7 covers the traps.

Only then consider the model

By this point you will know whether the model is the problem. Most of the time it is not, and a model switch is a distraction until the other stages are measured.

Where this course connects

This course is one part of a larger set. The links below take you to the courses and labs that go deeper on each area, and each one is worth reading next to the module it matches.

If you want the observability side in writing, the blog post on observability for AI agents covers what to track and how to set up the tooling.

Where to go from here

The model gets the blame because it is the only part of the system that talks back. Everything else is quiet until somebody times it. Time it first, and you will usually find that the slow part was never the chef.

Chapter summary
  • Every stage of a RAG answer has a module in this course, and every fix maps to one stage.
  • Stacking fixes one at a time shows which bar moves and why. The largest bar is usually the right next step.
  • No fix in the course changes how fast the model writes. The gains come from making the model start sooner and read less.
  • The method works on any bot: measure every stage, find the largest, change one thing, measure again.
  • The links in this module take you to the deeper courses and labs for each area.
Check your understanding
  1. Pip's team applied the queue fix and the index fix. The first word and the full answer both drop by the same 3.35 seconds. Why does the rest-of-answer span stay unchanged, and what would have to change to shorten it?
  2. Your bot has a 900 ms queue, a 300 ms search and a 1.1 second first word. Which stage would you change first, which module explains it, and what evidence would tell you the change is working?
  3. A colleague proposes switching to a faster model before any measurement. Using this course's method, describe the one measurement that would decide the question, and what result would make the switch the right call.

Finished this lesson?

Mark it done — your progress is saved automatically.