A year into evaluating Test, Learn, and Grow (TLG) our team at the GO Lab are learning as much about the problem of evaluating iterative, adaptive, and learning-oriented public sector reform as we are about the programme itself. This piece is a reflection on what that problem looks like from the inside.
When Test and Learn becomes the ‘thing’ being evaluated
The Magenta Book is the canon of government evaluation. Its new Test and Learn Annex is an odd addition because it does not describe an evaluation method. The Annex, as it makes clear itself, describes a way of developing and delivering policy by testing assumptions, learning from real-world evidence and adapting before interventions are taken to scale.
But it leaves a different question: What happens when the thing we want to evaluate is not only an intervention developed through Test and Learn, but Test and Learn itself as a way of working?
That question reflects a broader shift in how governments think about complex problems. As James Plunkett and others have argued, public policy rarely behaves like medicines. The hardest public service problems cannot be addressed following a linear design -> implement -> evaluate -> scale model. Interventions interact with institutions, professional practice and local context, and both problems and solutions can change through service delivery.
The Cabinet Office’s Test, Learn, and Grow programme, part of its wider public sector reform work, is an ambitious attempt to put this thinking into practice. It brings together central government with local practitioners in “Accelerators” working across different policy areas and places. Rather than beginning with predetermined interventions, Accelerators start with public service challenges, test assumptions and adapt what they do as evidence emerges.
The ambition goes beyond improvements for individual local interventions – this is the “Grow” in the name. Local teams often encounter constraints they cannot themselves remove: rules around data sharing, short-term funding, accountability and other features of the system whose levers of control sit elsewhere. Rather than treating places simply as testing grounds for policies developed in Whitehall, TLG aims to create a two-way-relationship between places and centre: learning from what happens locally, spreading useful practice, and addressing the conditions that get in the way of adaptive public services. In the programme’s own words, this is about creating a “new way of working in government” - in line with what the new government calls more ambitiously “Rewiring the State.”
Attendees at the TLG programme launch
What makes TLG difficult to evaluate
Three features of TLG make the evaluation unusually challenging.
First, both the “how” and the “what” can move over time. Accelerators are expected to adapt as evidence emerges, but learning can also change their understanding of the problem itself: what is preventing better outcomes, for whom, and therefore what should be ultimately tested. The programme aims not to simply improve the route to a fixed destination; the destination can move too.
This creates a moving-evaluand problem: if an intervention changes as evidence accumulates, what exactly is being evaluated? Its original design, its latest iteration, or the process through which it evolved? It creates an accountability problem alongside it, since the same trajectory can be understood as meaningful learning or as drift, as productive discovery or as delay. There is also the problem of causality: how do we establish whether changes in the how actually produce better interventions and better outcomes, rather than treating adaptation itself as evidence of success?
Second, there is substantial variation across Accelerators: they operate in very different policy settings, from children’s services to neighbourhood health and economic inactivity, and focus on different aspects of public service delivery, from digital and frontline services to inter-organisational collaboration. This variation is not implementation noise around a common intervention; it is inherent to the programme’s design, and therefore part of what the evaluation needs to understand rather than simply control for. How, then, do we identify patterns without treating fundamentally different interventions and contexts as equivalent, and what can findings from one Accelerator tell us about TLG as an approach rather than about place?
Third, there is scale. TLG’s ambition does not stop with better local interventions. Learning from places is intended to identify conditions that enable or constrain this adaptive way of working and ultimately contribute to changes in central government rules, routines, and relationships. The evaluative object therefore stretches from changes in particular services through changes in how teams work and interact with partners, to institutional change at the centre. How do we trace the connections between an Accelerator's work locally and changes in the wider conditions under which government works? Where along that chain can credible contribution claims be made, and where does the evidence run out?
Our initial evaluation design anticipated many of these challenges and was deliberately built to adapt, drawing on established traditions, such as developmental and realist evaluation. What we could not know in advance was what that adaptability would require in practice. Over the first year, we have had to reconsider where to concentrate evaluative effort, what can be realistically measured, which methods remain appropriate as the programme changes, and how the different strands might ultimately fit together.
No single vantage point
No single vantage point brings the full programme into view. The evaluation has gradually settled into a set of more bounded enquiries: choices about where to look closely, made in the knowledge that looking closely anywhere means not looking elsewhere.
What holds these vantage points together are that they answer to the same three questions: whether a test and learn approach produces better outcomes than business as usual, through what mechanisms, and at what cost. The questions are familiar, but the challenge is answering them when adaptation is an intended feature of the programme rather than simply a feature of its context.
Our in-depth qualitative fieldwork in a small number of Accelerators is where we can watch the evaluand move. It lets us walk alongside a problem being redefined as it happens and record what teams knew and when, rather than reconstructing it later. The trade-off is reach: we know a few places well, the rest of the portfolio much less well, and we cannot assume the places we chose stand in for the ones we did not.
Our longitudinal survey across all Accelerators and central government teams supports cross-accelerator analysis. Common measures repeated over time turn variation into something we can analyse rather than something to be controlled away, and let us ask whether a pattern found in one place is also true of others. However, what a survey cannot offer detailed explanations, and its reach depends on response rates we do not control.
The Grow strand of our evaluation asks whether constraints identified in places reach the centre, and what happens to them when they arrive. That means following particular problems forward rather than assembling a story afterwards from whatever changed.
Impact and value for money analysis face a different challenge and this is also where evaluations of adaptive programmes are most at risk of stopping short of aspiration. A causal design needs a stable object to attach itself to, and a costing needs to know the boundary of resource use; neither is available while an intervention is still taking shape. Our impact strand is using an evaluability matrix across sites to identify where credible causal analysis may become possible. Crucially, it also shows where it is not yet possible and highlights the necessary pre-conditions to enable such analysis. The value for money strand is doing the equivalent work by building the framework and data foundations needed for later economic assessment.
The different strands are looking at different parts of TLG, but ultimately their evidence has to come together around the larger question of: does a test-and-learn approach to policy making contribute to better outcomes? We must be humble and acknowledged that there is no substitutability here: observing adaptation is not evidence of better outcomes, while demonstrating impact for a particular intervention does not by itself show that Test and Learn produced an effect. What we are still trying to build is a robust architecture of inference: an account of how evidence generated at different levels, through different methods and over different timescales can legitimately be brought together to support that claim - and where it cannot.
Moving with the programme – but not too much
An evaluation of a dynamic programme like this is expected to be useful while the programme is still moving. Evidence that arrives after an Accelerator has moved on, or after a decision has been taken, may be rigorous but not very useful. Being close to delivery is what makes it possible to feed findings back while there is still scope for TLG teams to take action. But that closeness creates its own tensions.
The first is timing. Evidence does not mature at the speed of delivery. Early findings rest on limited observation, are often not yet triangulated, and sometimes look different six months later. One arrangement we have arrived at is a four-month reporting cycle: long enough to accumulate and analyse data properly, short enough that findings still meet a live decision. What the cycle cannot do is make an early finding more certain than it is, so the discipline is in how findings are stated rather than in when they are released.
The second is coherence. If every part of the evaluation moves with the programme, there is nothing left to measure change against. Some things are held still on purpose. Core longitudinal survey measures stay consistent across waves even as other parts of the evaluation respond to new questions. The overarching evaluation questions do not change to fit what the programme happens to produce, and neither do our standards for critical appraisal or interpretation of findings.
The third is independence. TLG itself involves a great deal of reflection and learning, so programme learning and evaluation do not happen in separate worlds. We compare interpretations, feed findings back, and sometimes influence what happens next. Analytical independence therefore rests less on separation from the programme than on being answerable for our own judgements the evaluation makes: the standards we hold evidence to and the findings we put our name to. For us, that responsibility is underpinned by the University’s research governance. This tension does not entirely disappear, though: when a finding we have fed back becomes part of what we later evaluate, it can be difficult to disentangle where the programme’s own learning ends and where evaluation begins.
What our first year suggests is that a highly dynamic, iterative, and adaptive programme requires neither an entirely adaptive evaluation, nor a fixed one. Different parts move at different speeds, and we need to make deliberate choices about which is which. Distance, both in time and in analysis, is less a principle we have arrived at than a setting we keep adjusting from inside the programme.
What cannot move
Distinguishing between what we did not find and what we were never positioned to see is one of the harder tasks in an evaluation like this. Learning may stay local. New ways of working may not produce better interventions. Better interventions may not show up as better outcomes during the evaluation period. A finding that the expected connections fail matters as much as a finding that they hold.
TLG is an experiment in how government develops and improves public services. A year in it has also become an experiment for us in how this kind of programme can be evaluated. Our evaluation design has not simply been applied to a changing programme; it has itself had to develop in response to what we have encountered. That means continuing to iterate parts of the evaluation architecture while keeping others stable, and exploring new forms of evidence where they may add value. But methodological adaptability and closeness to delivery cannot mean adapting the standard of evidence to fit what the programme happens to produce. A critical evaluation must preserve the possibility of saying that an expected mechanism is not visible, that evidence at one level does not support a claim at another, that a question cannot be answered credibly with the available data and timeframe, or that government has become more iterative and adaptive without this (yet) producing better outcomes for the people it serves.
Above all this sits a problem of accountability. In an adaptive programme, when does changing course count as learning, and when as drift? That is a question about how evidence is understood, not only about where and how it is gathered, and it is one the programme and its evaluators face together. Working out what a defensible answer looks like, and who is entitled to make it, is part of the difficult work for evaluation teams.
A year in, we do not have settled answers. The new Test and Learn annex to the Magenta Book sets out how evaluation can support adaptive policy development and how iterative testing can prepare interventions for robust evaluation. That work is already going on within TLG, supported by Cabinet Office evaluation colleagues and advisers working alongside Accelerators, which creates the potential for a wider evidence base. TLG raises a further question: what does an architecture of inference for adaptive programmes actually look like, and what can it credibly support?
Beyond producing findings about TLG, we hope this evaluation will contribute to that: to a more mature practice for evaluating adaptive government, one that is methodologically plural without becoming analytically baggy, willing to experiment without mistaking novelty for rigour, and explicit about both what the evidence allows us to know and what it does not.
Thanks to the GO Lab TLG evaluation team - Eleanor Carter, Ashna Devaprasad, Juliana Outes Velarde, Maria Patouna, and Ben Eyre - for reading drafts of this and pushing back where it needed it. The reflections, and any mistakes in them, are my own.