Steering a team through the shift to AI
TBC
velocity change
TBC
AI touchpoints per sprint
TBC
team confidence score
The ask: keep delivering nationally while rebuilding how the team works
As Design Lead for Health Choices & Prevention, I own the national third-party integrations landing in the NHS App, and for the past year I’ve also been rebuilding how the team behind them works.
The delivery model was built for yesterday.
We ran the way most mature teams run: sprint cadence, specialist roles, gate based assurance, a lot of ceremony. Human paced, and this was fine for the work it was designed for.
However, AI then changed what a designer could do in a day faster than the traditional process would allow. Exploring what the process looks like tomorrow was a must.
I’m rebuilding how the team works, not just what they use
I’m moving us to an AI-enabled model, and it looks different for each design craft.
For interaction, we stopped handing over Figma files. Designers now prototype AI-first in HTML, hosted on Heroku, so developers can inspect the real thing directly rather than rebuilding from a screenshot.
For content design, I worked with the designers and explored building a shared skill so writers can iterate faster and get the reasoning behind a change, not just the words.
For research, the goal is AI analysing session notes to surface patterns. It would save real time, but it’s also where bias creeps in most easily, so it has to be handled carefully. Client governance hasn’t approved it yet.
The result is a shared library of skills the whole team can draw on, and a process that adapts around where the handoffs actually need to be rather than where the ceremony put them.
So how do you measure AI?
Something everyone’s asking, and I don’t think anyone’s fully answered it.
For HC&P, I set a monthly loop against account-wide agreed metrics: velocity, AI touchpoints, and team confidence.
Right now I’m most interested in what the team says. Where they’re reaching for it, where we can work together, where we push it forward. We use what we can within the client’s guardrails, and we experiment. So confidence is the one I’m watching closest, because it tells me where the AI is actually reliable and where it isn’t yet.
Not a retrospective after the fact. A live loop that adjusts tooling, support and approach whilst delivering national patient health improvements.
Why I’m doing it in the open
This is a change in progress, and it’s not finished.
Rebuilding how a team works while it’s still delivering nationally is harder than doing either on its own. But a delivery model isn’t a thing you design once and hand over. You steer it, continuously, alongside the people doing the work. The steering is the job.