Partnership opportunities

Secure your seat

Call to action
Your text goes here. Insert your content, thoughts, or information in this space.
Button

Back to speakers

Dr Nikolay
Burlutskiy
Head of Applied AI
BBC
Nikolay is currently Head of Applied AI at the BBC, before that he was a Senior Engineering Manager leading enterprise AI platforms at Mars, where he led building and scaling of the Mars AI Experiences (MAX) platform - an AI capability deployed across a global organisation of 170,000+ associates. Previously Director of ML & AI at AstraZeneca (UK), where he led teams of AI scientists and engineers applying machine learning to accelerate R&D pipelines. Earlier, he led AI technical product development at ContextVision (Sweden), contributing to the creation of Inify, a spin-out AI company valued at $40M. Nikolay began his career as an AI engineer at Samsung Electronics (South Korea). His experience spans building AI systems end-to-end—from hands-on model development to leading large-scale AI programmes, shaping strategy, and delivering production systems in complex environments. PhD in AI with 35+ peer-reviewed publications and US patents. Active contributor to the AI community as a speaker, conference organiser, and reviewer for leading venues including NeurIPS, ECCV, and MICCAI. Focused on building and deploying trustworthy AI systems at scale that deliver measurable real-world impact.
Button
01 December 2026 11:00 - 11:30
The economics of inference: What generative AI actually costs in production
The token price of a frontier model has fallen roughly 70 percent since GPT-4's original 2023 launch price, yet blended per-token cost for the newest frontier models has actually risen around 100 percent since the start of 2026, even as mid-tier and budget models keep getting cheaper, down roughly 36 percent year over year. Inference has now overtaken cloud infrastructure to become the second-largest line item in enterprise AI budgets, trailing only talent spend. This session covers where that money actually goes and where it comes back, from model tier selection through to the specific engineering levers that move a real invoice. Key takeaways: - Why frontier and budget model pricing have split into two different markets, and what that means for which model tier belongs on which workload - The five levers that move inference cost the most: quantisation, response caching, prompt compression, model routing and batch processing, and the realistic savings range for each - Why a straight swap from a frontier proprietary model to an equivalent open-weight model can cut monthly API spend by 80 to 95 per cent for the right workload - How to build an inference budget that survives the next 12 months of pricing swings, not just the current one