We made Tempomat's AI agent faster by cutting model calls, not by making each call quicker. Across 26 realistic questions, the median answer went from 4 model calls to 2, from 84.5K to 40.4K input tokens, and from 16.9 to 11.8 seconds on the clock, while correct answers rose from 24 to 26.
Tempomat is an AI media-buying copilot. People ask it things like "how did TikTok do this month?" or "raise the budget on my best campaign", and it answers with charts, tables and approval cards built from the ad platforms' APIs. This is a case study of our own product, measured on our own machine; there are no customer numbers in it. All figures come from our internal performance write-up, comparing a baseline from September 27, 2026 with the version after the changes.
What was the problem?
Answers felt slow, and they were. The median question took about 17 seconds, and one in ten took more than 33. Store and TikTok questions were the worst: four or five model calls and up to 200K input tokens for what should be a single lookup. And two of our 26 test questions got wrong answers, including one that said a TikTok connection was revoked when it wasn't.
Our agent runs on eve, Vercel's open agent framework (in beta), with an OpenAI-compatible model behind it. A turn is a loop of model calls, which eve calls steps: the model reads the whole context, picks tools, the tools run, and the model reads everything again. Every step re-sends the context and waits for the provider.
Where does an AI agent's time actually go?
We read the traces before changing anything. The answer was clear: model calls dominate. A single call took 2–4 seconds in a sequential run and 7–14 seconds under eval concurrency, almost regardless of how much it wrote. In one trace, a 19-token reply took 6.5 seconds. So the biggest lever was not faster calls but fewer of them.
| Cost | What we saw | Size |
|---|---|---|
| Model calls | 2–4 s each sequentially, 7–14 s under concurrency, whatever the output length | Dominant |
| Tool search | 1–4 extra steps per platform answer; the found tools joined the tool list and broke the prompt cache (20K cached tokens → 0) | 1–4 steps + uncached context |
| Loading a skill | A whole step before any work, in almost every data answer | 1 step |
| Connection set-up | 7 availability checks re-run on every step | ~150 ms per step |
| Event logging | One awaited database write per event on the critical path | ~0.4–0.6 s per step |
The tool search finding deserves a sentence of its own. Prompt caching lets the provider reuse the unchanged start of the context between calls, which is cheaper and faster. Each search added newly found tools to the tool list, the start of the context changed, and the cache was gone. Every later step then grew from about 20K to 30–45K uncached tokens.
What did we change?
- Root tools for everyday reads. The common store and TikTok reports used to sit behind a tool search on each platform's MCP server. We exposed eight of them directly to the agent, with the same input schema and description, read-only, and blocked the duplicate names on the MCP side so there is one path per read. No search, no cache break.
- Batch every known call. The base instructions now tell the model to make every call it already knows it needs in one response, and to load a skill in the same response as its first data calls instead of in a step of its own. Skills stopped calling for a context lookup the session brief already contains.
- Defaults instead of questions. "Weekly report" now means the last complete week instead of a form asking which week, and date words map to fixed presets.
- Memoized connections. The availability of each platform connection is cached for 60 seconds per brand, instead of being re-queried before every step. The access token itself is still read fresh for every call.
- One database write per tool round. Events within a step are buffered and written together, instead of one awaited write each.
- Leaner tool schemas. The JSON schemas the model sees dropped long regex patterns for ids and dates (validation still happens in code), and our MCP route verifies only the token signature for the two request types that don't touch data.
Two accuracy fixes came out of the same traces. A brand memory saved by one test ("approvals are reviewed by our media buyer") was being read by other tests as a rule to ask for a review in text before the approval card. Now a memory can't add confirmation steps to an approval-gated action: the card is the review. And the TikTok answer is right again, because it reads the data directly instead of trusting a connection status.
What were the results?
| Metric | Before | After | Change |
|---|---|---|---|
| Correct answers | 24 / 26 | 26 / 26 | +2 |
| Wall time, median | 16.9 s | 11.8 s | −30% |
| Wall time, 90th percentile | 33.3 s | 27.4 s | −18% |
| Agent turn time, median | 13.8 s | 8.8 s | −36% |
| Model calls per answer, median | 4 | 2 | −50% |
| Model calls, all 26 prompts | 93 | 66 | −29% |
| Input tokens, median | 84.5K | 40.4K | −52% |
| Input tokens, 90th percentile | 166.4K | 79.5K | −52% |
| Tool searches / skill loads | 13 / 14 | 1 / 3 |
The full eval suite moved the same way. Passing evals went from 94 of 97 to 102 of 103. The median time per turn went from 18.2 to 16.4 seconds and the 90th percentile from 57.1 to 39.6 seconds; input tokens at the 90th percentile went from 121.6K to 72.9K.
The biggest single wins were exactly the answers that hurt most. Store, cash-on-delivery and TikTok questions went from 4–5 model calls and 95–200K tokens to 2 calls and about 40K tokens, answered in 9–10 seconds. The weekly report now saves in 3 steps instead of stopping to ask which week.
What didn't help?
- Lower reasoning effort for the final prose step. Back to back on fresh servers it made no measurable difference (median 2.5 s at the default vs 3.0 s at minimal), and an earlier pair showed the opposite. We removed the setting.
- A faster model for simple questions. Our fast and main model are currently the same, so there was nothing to measure.
- Ending the turn right after the chart. The framework has no "stop after this tool" option, so the short prose step stays. It costs about 2.5 seconds now that the context is mostly cached.
- Several model calls per durable step. It would widen what gets replayed after a crash, and not every platform write has an idempotency key. Safety beats a second saved.
What did we learn about making an AI assistant faster?
- Count steps first. Every model call avoided saves seconds; nothing else comes close.
- Protect the prompt cache. Anything that changes the start of the context, like adding tools mid-turn, makes every later step slower and pricier.
- Give common questions a direct path. A search is fine for rare tools, not for the five questions people ask every day.
- Default, don't ask. A sensible default with the period shown beats a clarifying question.
- Measure accuracy with speed. Our fastest version is also the most correct one, because several slow paths were also wrong paths.
The changes had to keep our safety rules intact: every campaign is still created paused, and anything that spends still parks on an approval card. How we check that on every change is in testing an AI media buyer. What the agent can do for your accounts is on the features page, and how it reaches Meta is in our Meta Ads MCP guide.
Why are AI agents slow?
Mostly because each answer is several model calls in a row, and each call re-sends the whole context. In our traces, model calls dominated everything else, regardless of how long the reply was.
How do you reduce AI agent latency?
Cut the number of model calls: batch tool calls into one response, give common questions a direct tool instead of a search, use defaults instead of clarifying questions, and keep the prompt cache intact.
Does prompt caching make agents faster?
Yes, when the start of the context stays the same between calls. Adding tools in the middle of a turn changed ours and threw the cache away, which made every later step larger.
Did making the agent faster make it less accurate?
No. Correct answers on our 26 test prompts went from 24 to 26, and passing evals from 94 of 97 to 102 of 103.
Tempomat answers from Meta, TikTok, Google Ads and your store in seconds, and asks before anything spends.
Start your 7-day free trial