Slow AI feels broken before it is wrong
Users forgive a loading spinner only when the payoff is obvious. Many AI features hide retrieval, tool calls, planning, generation, and validation behind one blank wait.
The useful signal is usually already there.
Hit the button and watch the hidden AI surface area light up.
Budget the wait
Choose the model, context size, streaming strategy, and tool count around the user moment. A background research job can take longer than an inline autocomplete.
Latency budgets turn performance into an explicit product decision.
Show progress with receipts
If the job is slow, expose useful progress: sources found, checks complete, draft ready. Streaming text can help, but structured progress is often calmer.
Flip the switches and watch the story turn into a tiny operating model.
Pick the right amount of intelligence
Latency is not just infrastructure. It is the shape of the promise you make to the user.