Optimize

Spend the fewest tokens for the best result

A complete, tool-by-tool playbook for token and credit optimization — from prompt caching and model routing to the viral Caveman method. Plan first, pick the smallest sufficient model, reuse outputs, and batch.

The four moves that win every time
Plan first
Architect & draft in a free chat before spending paid tokens/credits.
Smallest model
Route by difficulty; escalate only on low confidence.
Reuse outputs
Cache prefixes, cache answers, reuse renders & clips.
Batch
Queue non-urgent work for ~50% off.

Pricing, discounts, and TTLs change frequently. Always confirm current numbers in the provider's own docs before relying on them for budgeting. Last reviewed June 2026.