The Goldilocks Zone: Matching Goals to Interventions
Why the size of your goal must match the size of your intervention, what nudge meta-analyses really tell us about effect sizes, and how to balance exploration and exploitation when resources are scarce.
By Erik Bohjort, licensed psychologist and creator of the CLEAR Change Framework, Stockholm, Sweden · September 2026 · 6 pages, PDF
In brief
Between any group of people and any outcome sits a stack of barriers, and a goal is a statement about how many people must get through it. Ambitious goals require solving most of the barriers; modest goals can be met by solving a few. This paper by Erik Bohjort, licensed psychologist and creator of the CLEAR Change Framework shows that most behaviour-change programmes get the relationship wrong in one of two ways: they set transformational targets and fund a single nudge, or they build an elaborate intervention for an outcome a light touch would have delivered.
To calibrate the match the paper reviews the nudge meta-analyses. Mertens et al. (2022) reported a medium effect (Cohen's d = 0.43) with publication bias; Maier et al. (2022) reanalysed the same data and found a bias-corrected d between −0.01 and 0.08; DellaVigna and Linos (2022), using the complete record of 126 trials run by two US government nudge units covering 23 million people, found an average 1.4 percentage-point improvement against 8.7 points in comparable academic publications. The practical reading: a typical cheap, communication-style nudge moves a behaviour by a small single-digit number of points, and around two points is a sensible planning figure.
The paper then turns the logic around, letting available resources set the goal, and borrows the multi-armed bandit from decision theory to show how to split a budget between exploiting proven interventions and exploring higher-variance structural ones. A worked example takes a screening-attendance programme from a 60 percent baseline toward an 85 percent target by tiering the goal, and the closing section places the argument inside CLEAR's Clarify, Leverage and Experiment steps.
Key points
- High goals commit you to solving most of the hindrances; low goals permit you to solve a subset. No intervention design escapes this arithmetic.
- Calibration rule: a single light-touch intervention buys roughly a two-point improvement. A twenty-point goal needs ten independent interventions that stack, or a different class of intervention: defaults, structural redesign, or removing the behaviour.
- DellaVigna & Linos (2022): across every trial two US nudge units ran, the typical nudge added 1.4 percentage points on a baseline take-up of roughly 17 percent; publication bias and low power explained about 70 percent of the gap to published estimates.
- Two ways to miss the zone: too hot (a large goal on a small intervention, the common failure) and too cold (a small goal on a large intervention, rarer but costlier, and hidden because success hides waste).
- Three bandit principles: explore in proportion to the time horizon; explore in proportion to the gap between what you can prove and what you need; explore cheaply and often, not expensively and once.
- A unit that delivers a two-point improvement on twenty behaviours a year is worth more than one that promises twenty points on one behaviour and misses.
The Goldilocks Zone: Matching Goals to Interventions
Why the size of your goal must match the size of your intervention, what nudge meta-analyses really tell us about effect sizes, and how to balance exploration and exploitation when resources are scarce.
1. Goals and barriers are the same object seen from two sides
Between any group of people and any outcome sits a stack of barriers. They do not notice. They do not care. They cannot find the time. The tool is confusing. The manager disapproves. The habit is strong. The alternative is cheaper. Each barrier removes some fraction of the people who would otherwise reach the outcome.
A goal is a statement about how many people must get through the stack. Set the goal at 5 percent improvement and you may only need to remove the one barrier that is stopping the most persuadable group. Set it at 50 percent and you must remove nearly every barrier, because any one of them left standing will stop enough people to keep you short.
This sounds obvious written down. It is violated constantly in practice, because goals are set by one group (leadership, funders, strategy) and interventions are designed by another (delivery teams, agencies, behavioural units), and the two rarely sit down to check whether the barrier stack the second group is willing to attack is large enough for the target the first group has announced.
2. What a nudge is actually worth
To calibrate the match, we need honest numbers about what interventions deliver. The nudge literature provides the cleanest case because it has been meta-analysed several times with conflicting results, and the conflict itself is instructive.
| Source | Sample | Headline effect |
|---|---|---|
| Mertens et al. (2022), PNAS | 212 published studies, 447 effects | Cohen's d = 0.43, a "medium" effect, with moderate publication bias detected |
| Maier et al. (2022), PNAS, reanalysis | Same database, bias-corrected | d between −0.01 and 0.08 depending on method: "no evidence for nudging after adjusting for publication bias" |
| DellaVigna & Linos (2022), Econometrica | 126 trials, 23 million people, every trial run by two US government nudge units | Average 1.4 percentage-point improvement (about 8 percent relative), versus 8.7 points in comparable academic publications |
The DellaVigna and Linos result is the most useful because it is a complete record: every trial two nudge units ran, published or not. On a baseline take-up of roughly 17 percent, the typical nudge added 1.4 points. Publication bias and low statistical power explained about 70 percent of the gap between that figure and what appears in journals. The remainder came from differences in the interventions themselves: academic studies more often tested defaults and in-person interventions, while the units mostly tested letters, emails and reminders.
The practical reading is not "nudges do not work". It is that a typical, cheap, communication-style nudge, deployed at scale, moves a behaviour by a small single-digit number of percentage points, and that anyone promising more from that class of intervention is quoting a biased literature. Around two percentage points is a sensible planning figure. That is a real, cost-effective gain when the base is millions of tax filers or patients. It is nowhere near a transformation goal.
3. The two ways to miss the zone
Too hot: a large goal on a small intervention
This is the common failure. A board sets a target of doubling adoption, halving churn or cutting energy use by a third. The delivery team is given a communications budget and a quarter. They design a well-crafted nudge, run it, and report a statistically significant 2.3 percent improvement. Everyone is disappointed. The intervention was fine. The match was wrong. Worse, the disappointment is usually charged to behavioural science ("we tried nudging, it did not work") rather than to the mismatch, and the organisation loses a tool that would have been perfect for a goal of the right size.
Too cold: a small goal on a large intervention
The rarer but costlier failure. A modest outcome, such as getting a few more people to complete a form, is attacked with a redesigned process, a training programme, new incentives and a change-management workstream. The barrier stack for that outcome had one meaningful layer. Removing it with a default or a reminder would have done the job at a hundredth of the cost. The over-built intervention succeeds and is celebrated, and nobody notices the wasted resources because success hides waste.
4. Turn the logic around: let resources set the goal
Goal-setting usually runs from ambition to plan. The Goldilocks logic also runs in reverse, and the reverse direction is often more honest. If the resources available for a solution are small, the goal for its impact should be small. A team with one analyst, one email channel and no authority over systems should promise a two-point gain and deliver it, rather than promise a transformation and deliver a two-point gain.
This is not defeatism. Small, reliable, cheap gains compound. A unit that delivers a two-point improvement on twenty behaviours a year is worth more than one that promises a twenty-point improvement on one behaviour and misses. And a track record of matched promises is what earns the authority and budget needed to attempt the structural changes that deliver large effects. The path to big goals runs through correctly sized small ones.
5. Exploration and exploitation: allocating what you have
Once a goal and a budget are matched, the next question is how to spend the budget across approaches. Here decision theory offers a precise vocabulary. In a multi-armed bandit problem, a player faces several options with unknown payoffs and must decide, at each turn, whether to exploit the option that has performed best so far or explore a less-tested one that might be better. Exploit too early and you lock into a mediocre option. Explore too long and you spend the budget learning instead of earning.
Behavioural programmes face the same trade-off. Exploitation is deploying the intervention with the best evidence, typically a proven default or a well-tested reminder, and accepting its known, modest effect. Exploration is testing something new, such as a structural redesign, a novel framing or a removal of the behaviour, with a higher variance of outcomes and a chance of a much larger effect. Three principles from the bandit literature translate directly:
- Explore in proportion to the time horizon. A programme with one shot at a result should exploit. A programme that will run for years should explore heavily early, because a better arm found in year one pays off in every later year. Most organisations do the opposite: they explore in pilots that have no follow-on and exploit in long programmes that never test alternatives.
- Explore in proportion to the gap. If the best known intervention delivers two points and the goal needs twenty, exploitation cannot reach the goal, and the rational choice is to spend the budget looking for a higher-payoff arm. If the best known intervention already delivers the goal, exploration is a luxury.
- Explore cheaply and often, not expensively and once. Optimal bandit strategies test many arms with small samples before committing. In practice this means many small experiments at identified leverage points, run in parallel, rather than one large pilot of the option the team already preferred.
Risk appetite is what sets the balance. Exploration is risk-taking; exploitation is risk-avoidance. Neither is virtuous in itself. A programme that only exploits will never find the structural change that makes the goal reachable. A programme that only explores will never bank a result. The Goldilocks zone for resource allocation is a portfolio: enough exploitation to guarantee the modest goal, and enough exploration to have a credible chance at the ambitious one.
6. A worked example
A regional health provider wants to raise attendance at screening appointments. Baseline attendance is 60 percent. Leadership wants 85.
The barrier map shows six meaningful layers: patients do not receive the letter, do not read it, forget the date, cannot get time off, cannot get transport, and are anxious about the result. The evidence says an SMS reminder will move attendance by two to four points. That solves one layer. Reaching 85 requires solving at least four.
The Goldilocks response has three parts. First, restate the goal in tiers: 64 percent is achievable this quarter with reminders (exploit); 75 percent requires a default appointment slot with easy rescheduling and an employer letter (moderate exploration); 85 percent requires transport partnerships and a redesigned results process, and is a two-year goal. Second, allocate the budget accordingly: most of it to the proven reminder, a meaningful share to the two structural experiments, a small share to a genuinely novel arm such as mobile screening units that removes the travel behaviour entirely. Third, agree with leadership that the number they hear this quarter will be 64, and that it is a success.
7. Placing this inside CLEAR
The Clarify step of the CLEAR framework exists to set objectives and key results before any intervention is designed. The argument here sharpens what that step must produce: not just a goal, but a goal that has been checked against the barrier map (Leverage) and the planning effect sizes of the interventions the team can actually afford. The Experiment step is the exploration budget. The Analysis and Refinement steps are where the exploit-versus-explore balance gets updated with real data, exactly as a bandit algorithm updates its estimates after each pull.
8. Conclusion
Match the size of the goal to the size of the barrier stack you are willing to remove. Use honest planning numbers, about two points for a light-touch nudge, an order of magnitude more for defaults and structural change. Let small resources set small goals and be proud of meeting them. And when the gap between what you can prove and what you need is large, spend deliberately on exploration, because no amount of exploiting a two-point intervention will ever reach a twenty-point goal.
References
- DellaVigna, S., & Linos, E. (2022). RCTs to scale: Comprehensive evidence from two nudge units. Econometrica, 90(1), 81–116.
- Mertens, S., Herberz, M., Hahnel, U. J. J., & Brosch, T. (2022). The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domains. PNAS, 119(1).
- Maier, M., Bartoš, F., Stanley, T. D., Shanks, D. R., Harris, A. J. L., & Wagenmakers, E.-J. (2022). No evidence for nudging after adjusting for publication bias. PNAS, 119(31).
- Szaszi, B., et al. (2022). No reason to expect large and consistent effects of nudge interventions. PNAS, 119(31).
- Lattimore, T., & Szepesvári, C. (2020). Bandit Algorithms. Cambridge University Press.
- March, J. G. (1991). Exploration and exploitation in organizational learning. Organization Science, 2(1), 71–87.
© 2026 Erik Bohjort · clear-framework.com
Questions this paper answers
- How big an effect should I expect from a nudge?
- Plan on about two percentage points for a typical light-touch, communication-style nudge deployed at scale. That figure comes from the complete trial record of two US government nudge units (DellaVigna & Linos, 2022), which found an average 1.4-point gain, far below the 8.7 points reported in comparable academic studies. Defaults and structural redesign deliver an order of magnitude more.
- What is the Goldilocks zone in behaviour change?
- The range where the size of the goal matches the size of the barrier stack the intervention is willing to remove. Too hot is a transformational target funded with one nudge; too cold is a heavyweight programme for an outcome a reminder or default would have delivered.
- What does exploration versus exploitation mean for a behaviour-change budget?
- Exploitation is deploying the intervention with the best evidence and accepting its known modest effect. Exploration is testing something new, such as a structural redesign or removing the behaviour, with higher variance and a chance of a much larger effect. The right balance is a portfolio: enough exploitation to guarantee the modest goal and enough exploration to have a credible chance at the ambitious one.
More whitepapers
Attention Is the Scarce Resource
What shoppers who touch and look at products teach us about every behaviour we try to change, from checkout pages to shop floors.
Read the overviewDesign So the Behaviour Never Has to Happen
The most powerful behavioural solutions remove the need for behaviour altogether. Why defaults and structure beat persuasion, and how to use them.
Read the overviewCLEAR and the OECD’s LOGIC Framework
The OECD’s 2024 LOGIC principles tell institutions how to mainstream behavioural science. CLEAR tells a team how to run a change. How they compare, and where each needs the other.
Read the overviewApply this to a behaviour you need to change
Erik Bohjort works with organisations in Sweden, the Nordics and Europe to turn arguments like this one into measurable behaviour change.