Talk
Token-Maxxing Is Not a Strategy: Getting More Value per Token on Google Cloud
Per-token prices have fallen ~1000x in three years, yet in 2026, 73–93% of enterprises are over their AI budgets, and OpenAI's own data shows no link between token output and revenue. The industry calls this "token-maxxing": confusing activity with impact. More tokens are not more value. This talk flips the metric that matters. Instead of optimizing cost per token, we optimize reliability-adjusted cost per successful task, and see the counterintuitive result that spending more per call can make your system dramatically cheaper to run. Then we get practical, entirely on the Google stack. Live, we take an open-source Gemini model router and build a cost-aware layer. You'll leave able to cut your Gemini bill without dropping quality and knowing strategies to optimise your token usage.




