01 December 2026 11:00 - 11:30
The economics of inference: What generative AI actually costs in production
The token price of a frontier model has fallen roughly 70 percent since GPT-4's original 2023 launch price, yet blended per-token cost for the newest frontier models has actually risen around 100 percent since the start of 2026, even as mid-tier and budget models keep getting cheaper, down roughly 36 percent year over year. Inference has now overtaken cloud infrastructure to become the second-largest line item in enterprise AI budgets, trailing only talent spend.
This session covers where that money actually goes and where it comes back, from model tier selection through to the specific engineering levers that move a real invoice.
Key takeaways:
- Why frontier and budget model pricing have split into two different markets, and what that means for which model tier belongs on which workload
- The five levers that move inference cost the most: quantisation, response caching, prompt compression, model routing and batch processing, and the realistic savings range for each
- Why a straight swap from a frontier proprietary model to an equivalent open-weight model can cut monthly API spend by 80 to 95 per cent for the right workload
- How to build an inference budget that survives the next 12 months of pricing swings, not just the current one