As is already our habit, we like to step outside our own field of expertise every once in a while, and this month was no exception. We had the pleasure of hosting Fatos Ismali, an AI Architect at Microsoft with a background spanning Cloud Architecture, Data Engineering, and Deep Learning, for a guest lecture that tackled one of the most consequential shifts happening in software today: the move from applications to agents, and what it actually takes to bring AI systems into production.
The Shift: From Application to Agent
Fatos opened with a walk through the evolution of software itself – from classical computing, where abstractions stack upward in predictable layers, to the branching paradigms that neural networks introduced. He described how machine-learned systems represent a fundamentally different kind of “old machine,” one built on probability rather than fixed logic.
This distinction matters more than it might first appear. As Fatos explained, a prompt is a generative computational medium – there is no such thing as 100% accuracy, because everything is probability-based. This led to one of the more thought-provoking framings of the session: the difference between a model and an application. Models, he noted, don’t simply predict the next token – they predict the next state. That shift in thinking, from token-level prediction to state-level prediction, is part of what’s driving the industry from prompt engineering (2018/2019-era thinking) toward agentic engineering today.
Production Is the Real Bottleneck
One of the most direct points Fatos made was also one of the most practical: taking AI into production is the biggest blocker most organizations face. It’s one thing to experiment with AI, and another to actually operate it at scale.
He stressed that any organization trying to work with AI – or transition into working with AI – needs to seriously account for scale, capacity, and budget, three areas that are consistently underestimated. AI has real, and often substantial, costs behind it; Fatos gave a concrete example of AI infrastructure running €5 million a year, a figure that business owners frequently overlook when planning AI initiatives. The excitement around AI’s capabilities, in other words, doesn’t always come with a clear-eyed view of what it costs to run.
Questions from the Room
The session sparked genuine engagement from our team. One of our colleagues asked whether it’s possible to feed AI too much information, to the point where it actually becomes more confused rather than more capable at completing a task. Fatos responded by introducing the concept of context engineering – the discipline of shaping and managing what information a model actually receives, rather than assuming more input always leads to better output.
A follow up question was about how errors move through a chain of agents, prompting Fatos to touch on how failure can propagate across an agentic workflow – a reminder that as systems become more autonomous, understanding how they fail becomes just as important as understanding how they succeed.
Dea also left us with something to think about: if we’re writing emails with AI just for them to be read by another AI on the other end, where is this actually going? And if AI lets us do tasks faster, what are we really using that saved time for – more free time, or simply more tasks?

What the Data Says: The Stanford 2026 AI Index
To ground the conversation in evidence rather than speculation, Fatos shared findings from the Stanford 2026 AI Index Report, framing the discussion around opportunity versus risk. He walked through AI’s growing impact on the workforce, its influence on the software development lifecycle, and its expanding role in medicine and drug development – underscoring just how broad AI’s footprint has become across industries.
Beyond the numbers, he spoke to the tangible benefits already showing up in the form of workflow automation, and how organizations are starting to measure productivity gains not just at the individual level, but across entire teams – the difference between personal productivity and community productivity.
A Real-World Example: Vodafone UK
To bring the discussion out of theory, Fatos shared a case study from Microsoft’s work with Vodafone UK, illustrating how these concepts – agentic systems, context engineering, workflow automation – translate into measurable impact, particularly around reducing the time required to complete tasks.
He also touched on the importance of an AI gateway and pattern layer as a critical piece of infrastructure for organizations looking to deploy AI responsibly and at scale.

Key Takeaways from the Lecture
- Software is shifting from applications to agents, driven by a move from token-level to state-level prediction
- Production, not experimentation, is the hardest part of working with AI – scale, capacity, and budget are consistently underestimated
- Context engineering matters as much as prompt design – more information isn’t always better
- Real-world data, like the Stanford 2026 AI Index, shows AI’s growing impact across the workforce, software development, and medicine
- Case studies like Vodafone UK show that the benefits are measurable, not just theoretical
- Infrastructure choices, like an AI gateway, are becoming critical to deploying AI responsibly at scale
Sessions like this reinforce why we keep bringing outside voices into the room: the shift from applications to agents isn’t a distant trend, it’s already reshaping how teams like ours think about building, deploying, and scaling technology.
