In this episode, host Tessa Norman is joined by Leigh Bates, PwC’s Global Risk AI Leader, and Lilia Christofi, PwC’s EMEA Financial Services Data, AI and Tech Leader, to explore how financial services firms and regulators should respond as AI adoption accelerates and systems become more autonomous.
Against the backdrop of the FCA-commissioned Mills Review and the Financial Services AI Adoption Plan (an independent report published by HM Treasury), our guests unpack the shift from experimentation towards AI at scale. They discuss what greater autonomy could mean for firms and consumers, where existing regulatory frameworks may come under pressure, and how firms need to evolve governance and assurance as AI moves from supporting decisions to taking actions.
Finally, our guests share practical priorities for boards, and discuss the people, technology, data and operating-model foundations that firms need to be able to turn AI pilots into trusted, scalable deployment.
Listen on: Apple Podcasts Spotify
Tessa Norman: Hi everyone, and welcome to Risk and Regulation Rundown, the podcast where we explore the hot topic shaping financial services risk and regulation. I'm your host, Tessa Norman, and in this episode, we're exploring one of the biggest questions on the C-suite agenda: how should firms and regulators respond as AI adoption accelerates and as systems become more autonomous?
There’s been two major policy developments recently which have sharpened that debate. In early July, the FCA-commissioned Mills Review set out a seven-point agenda for preparing retail financial services for AI-enabled finance. Shortly after, the Treasury published the Financial Services AI Adoption Plan, an independent report which set out a series of recommendations to help the sector scale AI responsibly.
Today, we’ll be exploring what these reports tell us about how AI is changing financial services, and what firms and policymakers should do next in response. I’m really pleased to be joined by two brilliant guests to help us unpack all of this: Leigh Bates, PwC's Global Risk AI Leader, and Lilia Christofi, PwC's EMEA Financial Services, Data, AI and Tech Leader.
Welcome to the podcast.
Leigh Bates: Thanks, Tessa. Great to be here.
Lilia Christofi: Thank you so much.
Tessa: Leigh, let's start with your thoughts on that policy context that I set out in the introduction. The FCA has clearly been developing its approach to AI for some time, including through various testing initiatives. But it does feel like we're now moving into a bit of a new phase. How do the Mills Review and AI adoption plan build on that work to date? And what do you think those reports tell us about the future direction of travel?
Leigh: Well, July has been a busy month, hasn't it, in terms of those two reports being published recently. I think the clearest signal for me is continuity with much greater urgency in terms of what those reports talk to. Neither report argues for wholesale changes to that AI rule book, and no new AI rule book has been talked about as being the next step. But both recognise that the UK's technology-neutral, outcomes-focused framework is a real strength of ours. We have things like the Consumer Duty which tells us how to remain focused on customer outcomes. We have the Senior Managers Regime, which keeps accountability with named leaders. We have things like model risk, which provides disciplines for testing, validating, and monitoring systems, and then clearly operational resilience, an important topic to help firms understand how critical services could be disrupted and how they would continue to operate in the future.
What is changing now, and I think we would all agree with this, I'm sure we're going to build on this in the conversation, is the maturity of AI innovation now in the industry. We are now moving way beyond pilots and use cases. And until now, that's where much of the regulatory focus has been, particularly on experimentation and pilots, helping firms test use cases, understand the risks, experiment with new technology around interpreting rules. And the next phase for me is much more about making that support now work at scale. And firms now need practical, joined up guidance on how the existing frameworks apply when AI is embedded more into core operational processes and where changes occur more frequently as these new frontier models get updated on a regular basis and are getting more advanced. And when AI begins to take action, we've got to talk about agents. So, when AI takes action, that changes the risk profile as well in terms of what we need to consider. But the two reports I definitely see as being complementary. The AI adoption plan is much more focused on what needs to happen now to unlock responsible scaling. The Mills Review looks further out and looks at how firms and the financial services system and the regulatory framework may need to evolve as AI moves from supporting decisions to taking actions with agreed boundaries. But together, they're very complementary in terms of nature.
Tessa: There's a lot in those two reports and they cover quite a broad agenda. What are some of the messages that really stand out most strongly to you, Leigh, and how do they resonate with what you're seeing among firms as they move towards wider adoption of AI?
Leigh: There's three things that really stand out for me. The first is that the debate now is much more about value. So it's no longer about the number of pilots, the number of use cases, very focused on productivity and so on. That just needs to happen as table stakes in my mind. In other words, we are measuring outcomes of AI systems, not the outputs of AI systems. I think that's an important point to really think through in terms of where we are focusing our efforts. Now at PwC, we recently published our global AI Performance Study, which was a cross-sector piece of research, and that found that 20% of organisations we surveyed captured 74% of AI-driven returns, and the leading companies achieved over seven times the AI-driven performance of the rest. And the key difference was not that they used more AI, it was much more about they were making deliberate choices to focus on important business processes, business outcomes, and built the right foundations embedded into core AI workflows. Strongly supported, I would say, by a clear people agenda and the culture and change process to really drive adoption. The second is this point about autonomy. And this is about where AI is moving beyond summarising and drafting. And the key questions there are around accuracy, grounding, and appropriate reliance. When AI begins to plan and recommend, initiate processes, execute actions, the questions now change to permissions, accountability, consent, resilience, redress, and how the system is controlled while it's operating. That's an important distinction.
The third area I would say is governance. And I think Lilia and I would both agree on this point that governance is becoming much more part of the enabling infrastructure. The architecture point is an important point here. And policy generally expresses things like intent, but architecture really operationalises it. So, we have to think much more about things like policy as code and controls embedded into platforms and orchestration of workflows. And that really moves us from policy and standards to much more about operationalising responsible AI in practice.
Tessa: Lilia, building on that, what are your clients most focused on at the moment? And are there any aspects of the regulatory frameworks, whether that's in the UK or globally, that they're finding particularly either enabling or restrictive?
Lilia: We've always been talking about the fact that interpretation is quite difficult when you're looking at a cross-functional view, but even cross-institutional view across the financial industry. And it's important for us to make sure, as Leigh said, to codify controls, codify policies, codify regulation, and then know, based on the context and the intent of an agent, as we said, multi-agent systems are the thing of now and what we are building, knowing when to pull what without making the architecture really heavy, latent, and expensive. And that is quite difficult to achieve because to date what we have done is we've coded all the guardrails within agents or we've got control agents to do a lot of that work. And when you're looking at that at scale, that's not really possible, sustainable, operationally or otherwise. And so, one of the things that we're working on right now is how do we use ontologies, which is basically the institutional knowledge conversion into a codified, much more discreet set of rules that are going to be applied basically across the whole organisation. And understanding the traversal of using that knowledge, it's like being able to walk into a library and knowing exactly which book and where which shelf to get that book from based on the assignment that you have been given. And by creating that, you get repeatable outcomes that are much more deterministic because what AI is really good at is being probabilistic model. So, coupling that makes you a lot more assured as to the result that you're going to get and whether you are or not compliant.
So back to the whole dilemma. It's not about having good governance. It's about having governance in time, real time. It's about having control and evidence to support that real time, in time. And then it's about having the right operating model to support that being executed in the right way, which means where does the control lie within the whole scheme of the work? And more complex than that, if you think that 50% of the internet is agents, and agent-to-agent communication is happening, the biggest challenge that we have right now is where does responsibility lie? Because you could be compliant and ethical, but if you have another financial advisor that's not part of the institution, speaking to your institution and then speaking to the consumer, where does the consumer duty lie? So, it's an interesting dynamic we are having to work through.
Tessa: One of the concepts that stood out most to me from the Mills Review was the autonomy spectrum that it set out. It talks about AI taking on progressively more responsibility and the human role changing alongside that. You mentioned 50% of the internet is agents. Where would you characterise financial services as being on that spectrum at the moment? And can you share with us your thoughts on what is that greater AI autonomy likely to mean for both firms and consumers over time?
Lilia: My observation and I work with quite a lot of institutions across EMEA is that we're not quite ready yet across these institutions to have financial advisory type of agents available for people outside of the wealth spectrum and outside of PE-sponsored institutions. So, what you find is that you are going to very private or personalised or even your own codified financial advisory type agents or construct or using enterprise tools like Gemini, ChatGPT, open AI models and so on, on your own. And so, the question there that you have is even though they can make it transparent what the sources are of information that you are getting the guidance from. You as an individual, and we know that by looking at the consumer market, in effect, they don't all have the accessibility of the knowledge that they need to be able to validate whether the advice that they're getting is correct or not. And that's the biggest challenge, is who validates that institutional knowledge or that financial services knowledge whereby we make decisions from, whether you're a consumer or not. And unfortunately, from a maturity point of view, it will take quite some time for these tools to be available, but also to be integrated and procured directly through the frontier models. Because at the moment, there isn't that interconnect happening. It has to happen with regulators, and it has to happen with institutions themselves, not just relying on web scraping and unstructured data consumption.
Tessa: What are some of the key considerations for firms then if they're thinking about greater utilisation of AI agents? How would they need to gain confidence that the outputs and actions of those agents are accurate, trustworthy?
Lilia: Leigh sort of alluded to it in the earlier segment. When we used to code platforms, we always talked about the software delivery life cycle, and that was really part of defining the stage gates by which we feel comfortable that the code that we're progressing, we're happy with. In the new world, you have agent ops and you also have enterprise delivery lifecycle. So what you're expecting to happen is that you define your business requirements and all the specification mining, including the value engineering part of it, incorporated in as part of code repositories, which means your business users are now becoming a lot more technical and they're involved in the process and have access to inputting that data, be it unstructured, to be consumed into a data stewarded process in order to make sure that you can then validate that the output of the AI-codified solution actually meets those specifications and the ROI that has been defined within those documents.
As you progress those stage gates, you get to a point where you can create business scenarios, both negative, positive, with various paraphrasing and variations, which is critical because paraphrasing is very common amongst people, but we tend to validate the information or the question that we are getting asked. And so you have to build that as part of the test harness as you're progressing the code. So now you have various variations of outputs, and you are then looking at the spectrum to see how they cluster. If they cluster around the same sort of answer, you can have a certain level of confidence. And confidence is calculated through an algorithmic mathematical formula, which is evidence confidence, clarity confidence, structural confidence. There are various factors that we look at to be able to come up with that equation. And if that equation is above, let's say, 95% threshold or whatever the threshold you design, you can pass those stage gates to what we call the golden blueprint, and it's ready for deployment.
Tessa: So, we might think about how we would apply existing regulatory frameworks to this new and evolving world of more agentic, more autonomous AI. I can certainly start to think about ways in which that regulatory framework, if we think about governance, accountability, consumer outcomes, that might start to come under pressure. And indeed, that was one of the Mills Review's conclusions. Leigh, what are your thoughts on whether we might see that regulatory framework start to change, and how should firms as well be thinking about their own governance and controls and the way in which they might need to adapt them?
Leigh: Lilia has talked a lot about this already. I think just maybe to build on what Lilia has said, for me the pressure is less about whether existing obligations apply today, and more about how firms demonstrate compliance and trust in the future and become more dynamic and autonomous in the way they deploy AI into their environments. And lots of firms are asking us what good evidence looks like. How do they really think about the measurement you talked about, the confidence measurement, and that all comes down to how do we demonstrate consumer duty outcomes and personalise customer journeys? How do senior managers evidence reasonable steps when decisions are made across many different agents that may be operating in an orchestrated workflow? How do model risk and resilience really get evidence through third-party frameworks, particularly when we've got reliance upon third-party frontier model providers, which can change terms and conditions or change the models relatively quickly and we have to deprecate models and bring new ones in.
So, there's lots of points of change here that we need to think about how we best operate governance across all those things. But for me, the principle is that governance shouldn't really follow the label that's attached to the technology, it’s what the system can do. What is the functionality of the system in the first place? And I think we should be looking at assessing AI systems through four main lenses, I would say. The first is the level of autonomy. Second is the materiality and the reversibility of outcomes. How do we reverse out if something happens and something goes wrong? And its potential impact on customers. And then lastly, the firm's ability to reconstruct what happened. How do we intervene and provide redress? If we can recreate the process to understand exactly what's gone wrong. And that means moving much more beyond that point around human in the loop. We need to be much more specific in terms of what the human is doing, because otherwise it doesn't scale. You can't have a human in the loop across all workflows. We have to be clear in terms of what we're expecting humans to do, what information they receive, what actions they should take, when they should intervene, what escalates to them. And a lot of this comes down to the kind of runtime components, the operational runtime of AI deployments, which we tend to call AI observability. And this is a really important point around the future of monitoring AI systems in the future.
Lilia: Can I just add to that? It's very interesting because the more we build and the more that we put our institutional knowledge in, and it's cross-functional, we become more reliant, and reliance is actually quite a scary thing because once you're reliant, reversing back to paper is very hard. And a funny thing that I've observed looking at some of our clients and speaking to them, is that the staff that get accustomed and have adopted this profoundly, when it goes down, they would rather just wait for it to come back up again. That's where the resilience piece and simulating resilience is important.
Tessa: One of the other strong themes that came through in the reports, which links very closely to resilience, is this concept of sovereignty, and they talked about firms' reliance on a relatively small number of global AI and cloud providers. Lilia, what does sovereignty mean in practice for firms? How are they thinking about that? And what are some of the practical choices and challenges it might be creating for your clients?
Lilia: It's a hard topic. We talk about it and we just think that if we move our data and we put it somewhere that is local, then it's all going to be fine, it's all sovereign. But that's not necessarily the case. In the majority of the institutions that we work in, they're using SAS products. So, they're already reliant on third-party, have third-party risk exposure. A lot of these third parties are American offshore companies or otherwise. And so then from a human resources point of view, even from a delivery point of view, it's not onshore. Your exposure of your data and where your data is being used and how it's being used is never going to be onshore. But one of the concepts that we are trying to work with, especially in the European Union, because it is becoming more and more of a factor, across Europe, is that could we be running models and facilitating those through data centres, for example, that are open source, that are compressed and still deliver the level of quality that we require. And the answer to that is, I'm actually unsure, to be honest, because the complexity of the way that we are building the architectures at the moment and the reliance on the large language models is significant, which is why building that institutional backbone in terms of the knowledge graphs and the ontologies, and these are very technical terms, but they are very important concepts to be understood, are really important because they lessen the dependency on the language model. And what we are likely to find over time is that similar to what I've heard in industry through our alliances, is that organisations are going to invest more and more in building small language models of their own that they're going to be running. And so that is something that’s still sort of early days on, but probably in the next year or so, we'll probably see a change. And then the question is, who and how is that going to be controlled and governed? And how is that going to be treated in terms of IP? And it becomes ever so more complex.
Tessa: Leigh, it would be great to get your perspective on that governance and accountability point there.
Leigh: Look, the accountability point is probably the one for me that stands out in this topic. I think the principle for me is that a firm can outsource capability, but it can't outsource accountability. So, I think that's an important point. And we talked a bit about governance earlier. For me, the control object, if you like, is no longer the model itself. It's the entire system. It's the workflow end to end. And that's a distinction between where we've moved from things like model risk in the past, where we've been really focused on putting the right controls around the model when it was more of a machine learning deterministic model. Now, Lilia touched on this, we have to really think about data, clearly where data is accessed, where it's stored, the retrieval design, the orchestration design, what tools we call and where are those tools, what permissions that we hold in terms of who accesses what, as well as those human intervention points that we talked about earlier. So, I think sovereignty, and as we said, it's a complex topic and we could probably do at least an hour's podcast just on that topic. But it should be framed as maintaining this optionality point and resilience and control over critical data and compute and AI capabilities and not binary choices between domestic and international technology. It's much broader than that.
Tessa: We've talked a lot about a range of different issues, which are giving an important sense of the direction of travel, both in terms of how the industry is developing and in terms of the policymaking level and what policymakers and regulators are going to have to contend with. It'd be great to turn now to both your perspectives in terms of what firms can do today to prepare for that future and the way in which that's changing. Leigh, I'll come to you first. What should boards and senior management be prioritising in order to scale AI while maintaining trust and control?
Leigh: There's a number of different priorities, I would say. Maybe let me start with the top five that I can think of so far. I'm sure we'll build on this because there are many, many more than five. But I would say the first is knowing what you have today. And what I mean by that is building an inventory of current and planned AI systems including what you're doing with embedded third-party AI capabilities in your vendor landscape, including things like user-developed, no-code, low-code systems, and the more advanced AI agents. Now, that could include continuous scanning to discover AI systems that are being built across the organisations and scanning code repositories. But knowing what you have in place today is a super important point. Secondly, for me it comes down to risk appetite. So effectively, clearly defining what AI is allowed to do and what it isn't and where you are happy to deploy AI and setting the boundaries for data access for tool use, what has customer impact.
Third, is the point around governance. And this comes back to what I was saying earlier about industrialising and embedding governance away from static policies and standards to more consistent risk assessment processes, reusable control patterns, approved platforms and an automation of how you get that route to live through tooling, including the use of AI in governance. Fourth, I would say, it's the point that we talked about earlier about runtime controls. I'm really thinking about this point on, you need the right management information. All the things that Lilia was just talking about earlier around the knowledge graphs, the ontology, we need to service that up with a business lens on it so we can really understand how these AI systems are operating in production. And then the last point, I would say is a particularly hot topic right now and one that's very close to me, which is around trust and connecting trust to value. Our PwC research suggests that the leaders that are doing very well in this space systematically track business outcomes and make explicit decisions of what they're going to stop, what they're going to start, where they're going to scale. And really thinking about bringing the trust lens to that is a super important point. So that's what I would say are my top five.
Tessa: Brilliant, thank you. And Lilia, what are your perspectives on key actions for firms, perhaps in terms of, what are the key tech, data and operating model foundations that they need to be able to put some of those priorities that Leigh's spoken to into practice?
Lilia: We have to start with people, not technology and tools. And the people aspect is how do we align our operating model to work and surface the right information and the right design. So having your CISO, your CRO, your compliance team, your technology team, and then your data team working cohesively together. With bringing in different lines of business and really building for the enterprise is what's important, and what's missing, and I keep repeating this, whoever listens to me, is we are missing AI enterprise architects, and we need that enterprise architecture view because it starts connecting the dots on where the business strategically wants to head to, and then how do we enable the technology to be able to really serve that purpose? So, we're not building technology for technology reasons. We're building technology that is really helping our business be more adaptable, more able to be robust and flex, and agility in a technology architecture is what we always strive towards. So, we came out of plug and play architectures for years. We can't all of a sudden go backwards just because we're bringing in a new AI technology component in there.
We’ve talked about observability - it's funny because observability is actually management. And I think we tend to forget that. We're not changing the management aspect. We are managing non-humans. And the question is, how are we going to evaluate those non-humans against humans? And how do we build that into the architecture? So, thinking about those blueprints of what we put into the logic, which is management logic into the design, is actually important to be understood. And not everyone's a manager. So how are you going to create a task force that, A) is more technically capable, and B) able to manage? Well, if you think agents can't get angry, you've got another thing coming. And then I think, we talk about this all the time, and we have for decades, but data stewardship and data ownership doesn't disappear just because we put an AI system on top. In fact, now AI systems require AI stewardship and ownership. And I don't see a lot of that in industry. It's like, oh, well, we've got this agent. Who owns it? Who's the product owner? No one can answer. It does this thing. Okay, how can I reuse it? Does that change the ownership of it if I need to reuse it in a different way? We have to create those frameworks to begin doing the real work.
Tessa: Brilliant. Thank you both so much, and lots of food for thought and questions for firms to be asking themselves there. And to our listeners, if you'd like to find out more about any of the points that we've discussed today and some of the reports that we've mentioned, please do get in touch. If you enjoyed this episode, please do consider subscribing and leaving a rating or review, as it helps other listeners to find us. And finally, I want to share that this will be my last episode hosting this podcast series, but I'll be handing over the reins to my fantastic colleague Hinna Akhtar, so you'll be in great hands. Thank you.