AI by Hand ✍️

AI by Hand ✍️

Output Bound

Same tokens, different clocks

Prof. Tom Yeh's avatar
Prof. Tom Yeh
May 28, 2026
∙ Paid

Library › Token Problems

  1. Tokenization

  2. Subword Tokens

  3. Non-Word Tokens

  4. Token Ratio

  5. Headroom

  6. Token Pricing

  7. Prompt and Completion

  8. The Billing Line

  9. Asymmetric Pricing

  10. Words to Cents

  11. Token Mix

  12. Usage Forecast

  13. Standing Instructions

  14. Fixed Overhead

  15. Budget Ceiling

  16. Token Allowance

  17. Two Speeds

  18. Output Bound

  19. Rate Cards

  20. Usage Audit

Can two calls with the same total tokens finish at very different times? Yes. Input is read in one parallel pass; output is written one token at a time. A task that reads a lot and writes a little will almost always finish faster than one that reads a little and writes a lot. Agents designed to summarize or classify can be significantly faster than agents designed to generate long replies, even at the same token budget.

Paid members: the worksheet and its printable PDF are below ↓

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Tom Yeh · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture