AI performance

Model latency is a product choice

A smarter answer that arrives after the user has left is not smarter in product terms.

~5 minAI
scroll
01 · the shape

Slow AI feels broken before it is wrong

Users forgive a loading spinner only when the payoff is obvious. Many AI features hide retrieval, tool calls, planning, generation, and validation behind one blank wait.

trace · assistant response

The useful signal is usually already there.

guardrails armed

Hit the button and watch the hidden AI surface area light up.

02 · the move

Budget the wait

Choose the model, context size, streaming strategy, and tool count around the user moment. A background research job can take longer than an inline autocomplete.

latency-budget.json

Latency budgets turn performance into an explicit product decision.

03 · the check

Show progress with receipts

If the job is slow, expose useful progress: sources found, checks complete, draft ready. Streaming text can help, but structured progress is often calmer.

AI launch switchboard
sloppy
shippable0
fragile: the model is still freelancing without enough checks

Flip the switches and watch the story turn into a tiny operating model.

0milliseconds feels instant for tiny assists
0seconds needs visible progress
0background path for long jobs
04 · the takeaway

Pick the right amount of intelligence

Latency is not just infrastructure. It is the shape of the promise you make to the user.

source trail