Content library

The full TokenOps reference library

52 long-form artifacts — every guide, playbook, checklist, reference, and operating template from the TokenOps content pack. Read in the browser or download the source markdown.

Cost Anomaly Detection

Detect, investigate, and respond to unexpected token spend spikes before they break budgets.

LLM Gateway Architecture

Design a centralized gateway for tagging, routing, rate limiting, caching, and cost tracking.

Multi-Provider Routing Config

YAML reference for routing requests across OpenAI, Anthropic, Google, and open-source models.

Optimization Patterns — Detailed

Ten battle-tested patterns for reducing token cost without sacrificing quality.

Prompt Versioning Workflow

Treat prompts as production artifacts with version control, quality gates, and rollback.

Quarterly Cost Optimization Audit

Audit checklist covering tagging coverage, model mix, retry rates, caching, and waste reduction.

Model Upgrade / Migration Checklist

Validate quality, cost, latency, and rollback plan before swapping models in production.

Pre-Production Launch Checklist

Tagging, budgets, alerts, fallbacks, and incident readiness before launching an LLM feature.

Vendor Contract Negotiation

Pricing tiers, rate limits, SLAs, data-use, and termination terms to negotiate with LLM vendors.

TokenOps Case Studies

Composite optimization scenarios with starting state, interventions, and cost math. Company names and figures are illustrative teaching constructs, not measured production data.

Team RACI Matrix

Who is Responsible, Accountable, Consulted, and Informed for every TokenOps activity.

TokenOps FAQ

Answers to the most common TokenOps questions — from "what is it?" to multi-tenant billing.

TokenOps Glossary

Reference for every term used across TokenOps practices, dashboards, and contracts.

TokenOps Maturity Model

Score your organization across five maturity levels and plot a path to the next one.

Executive Briefing Template

C-suite-ready deck outline for presenting TokenOps results, asks, and roadmap.

Single-Model → Multi-Model Migration

8-week playbook to migrate from a single LLM to a routed multi-model architecture.

Multi-Tenant Billing Guide

Architecture and accounting patterns for charging tenants accurately for LLM usage.

RAG Pipeline Cost Optimization

6-week playbook to cut RAG token cost via chunking, reranking, caching, and retrieval tuning.

LLM Provider Comparison Matrix

Auto-generated from the maintained pricing dataset: per-model rates, cached-input pricing, context windows, batch discounts, and caching behavior with source links.

TokenOps KPI Dashboard Spec

Technical spec for a complete TokenOps monitoring dashboard — metrics, panels, alerts.

TokenOps Metrics Reference

Every metric used in TokenOps — definition, formula, owner, alert thresholds.

Tool Landscape Guide

Curated guide to gateways, observability platforms, and FinOps tooling for LLMs.

Incident Response Runbook

Step-by-step runbook for triaging and resolving token-cost incidents at any severity.

Quarterly Business Review

QBR template covering spend, savings, optimization backlog, and program asks.

ROI Justification Template

Build a defensible ROI case for funding a TokenOps program — costs, savings, payback.

SLA / SLO Definitions

Define and measure latency, availability, and quality SLOs for LLM-powered features.

Vendor Evaluation Scorecard

Scorecard for evaluating LLM API vendors across price, quality, security, and ops.

The 2026 LLM Pricing Landscape

GPT-5, Claude Opus 4.5, Gemini 3, DeepSeek V3.2 — per-1M-token pricing snapshot plus the tokenizer inflation trap.

Prompt Caching in 2026: The 90% Discount

Discount rates, TTLs, write premiums, and the break-even model for OpenAI, Anthropic, Gemini, and DeepSeek caching.

Reasoning Token Governance

How to right-size thinking budgets across o3/o4, GPT-5 reasoning, Claude Adaptive Thinking, and Gemini Deep Think.

The Loop Tax: Agentic Workflow Cost Patterns

The $47K runaway agent, the 5–30× estimation error, MCP schema tax, and the six required controls for agentic spend.

GreenOps & FOCUS: TokenOps' Environmental Twin

FOCUS 1.4/1.5 for AI, GreenOps research, and the dual-reporting playbook for dollars and carbon.

Enterprise TokenOps Case Studies (2025–2026)

AT&T 90% cut, fintech 73% saved, SaaS $48K→$19K — the stack that produced each result.

Semantic Caching in 2026: Beyond Exact-Match

The two-cache stack, similarity thresholds, invalidation strategy, and vendors — plus when semantic caching adds cost instead of saving it.

Small Language Models in Production (2026)

The 2026 SLM shortlist (Phi-4, Llama 3.3, Gemma 3, Qwen 2.5, Ministral, GPT-5 Nano), hosting break-evens, and the routing pattern behind 60–90% traffic shifts.

Multi-Agent Orchestration: The 2026 Cost Reality

Why multi-agent systems cost 5–15× single-agent equivalents, the context-inheritance tax, and the four cost controls that actually work.

Batch API Arbitrage: The 50% Discount You're Not Using

OpenAI, Anthropic, Gemini, DeepSeek, and Bedrock batch tiers — the migration pattern, the hidden traps, and the sharding trick.

Embedding & Vector DB Cost Tuning (2026)

Matryoshka truncation, chunking recall vs cost, reranking economics, and the object-storage vector-DB shift.

Structured Output & JSON Mode: Cost, Quality, Failure Modes (2026)

Why strict schemas add tokens, when JSON mode wins, the retry-loop trap, and the 2026 default pattern.

The TokenOps Observability Stack (2026)

The five categories (tracing, gateway, eval, FinOps, prod), the must-have span fields, and the five alerts that catch 80% of incidents.

Context Engineering: The Discipline (2026)

The seven context layers, per-layer token budgets, the five compression techniques, and the emerging Context Engineer role.

Fine-Tuning vs Prompting Economics (2026)

When to distil, when to prompt-cache, when to stay stateless — with break-even math and the managed-OSS-SLM shift.

TokenOps KPI Benchmarks (2026)

Four tiers of KPIs (cost, governance, quality, environmental), 2026 benchmark numbers, and the six questions a Level 4 program answers on demand.

Cross-Provider Arbitrage Playbook (2026)

The current arbitrage matrix, four arbitrage patterns, negotiation floor discounts, and cost-aware gateway fallback routing.

Prompt Compression Techniques (2026)

The compression ladder, instruction hygiene, RAG pruning, and how to gate compression on a golden eval set.

Output Token Control (2026)

max_tokens discipline, artifact-not-essay prompting, reasoning budgets, structured output economics, and retry caps.

Cache-Aware Prompt Architecture

Layered prompt layout, break-even math, TTL strategy, multi-tenant namespacing, and prefix-hash instrumentation.

Model Routing & Cascades

Static, cascade, and learned routing; task-to-tier mapping, escalation signals, and the quality gates before a swap.

Agentic Loop Cost Controls

Six hard controls, context-inheritance tax, tool schema tax, planning ratio, and the spans you must emit per step.

RAG & Retrieval Efficiency

Chunking that pays, retrieve-wide-send-narrow, Matryoshka truncation, query-side savings, and the metrics that matter.

Batch & Async Arbitrage

Qualifying workloads, stacked discounts, sharding and idempotency, hidden traps, and the pre-generation patterns.

The TokenOps Continuous Improvement Loop

Weekly, monthly, and quarterly cadences, the savings ledger schema, six drift alerts, and the culture rules that make it stick.