AI by Hand ✍️

AI by Hand ✍️

Two Speeds

Reading fast, writing slow

Prof. Tom Yeh's avatar
Prof. Tom Yeh
May 28, 2026
∙ Paid

Library › Token Problems

  1. Tokenization

  2. Subword Tokens

  3. Non-Word Tokens

  4. Token Ratio

  5. Headroom

  6. Token Pricing

  7. Prompt and Completion

  8. The Billing Line

  9. Asymmetric Pricing

  10. Words to Cents

  11. Token Mix

  12. Usage Forecast

  13. Standing Instructions

  14. Fixed Overhead

  15. Budget Ceiling

  16. Token Allowance

  17. Two Speeds

  18. Output Bound

  19. Rate Cards

  20. Usage Audit

Why does a short reply sometimes take longer than reading a large document? A call's time has two parts. Input is processed in parallel, every token in one pass, so even a large input clears quickly. Output is generated one token at a time, each waiting on the one before. For agents, this means latency is driven almost entirely by how much the agent writes, not by how much it reads.

Paid members: the worksheet and its printable PDF are below ↓

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Tom Yeh · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture