Can two calls with the same total tokens finish at very different times? Yes. Input is read in one parallel pass; output is written one token at a time. A task that reads a lot and writes a little will almost always finish faster than one that reads a little and writes a lot. Agents designed to summarize or classify can be significantly faster than agents designed to generate long replies, even at the same token budget.
Paid members: the worksheet and its printable PDF are below ↓

