top of page
Library - Portrait.png

Pilot Programs Should Test Decisions, Not Postpone Them

Aug 28
6 min read

Updated: Sep 2

BlindSpot Insights Strategy and Improvement branded article cover.

Pilot programs are meant to make decisions safer. They allow an organisation to test an idea at limited scale, learn what happens in practice and decide whether to expand, adapt or stop. Yet many pilots are launched without a precise decision in mind. They produce activity, meetings and encouraging anecdotes, but no agreed point at which the evidence must change what the organisation does.


That is how a temporary test becomes a holding pattern. The pilot is described as successful because it ran, participants were broadly positive or no major incident occurred. Meanwhile, the harder questions remain unanswered: did the approach improve the outcome that mattered, would it work under normal operating pressure, what would scaling cost and who is authorised to make the final call? A useful pilot does not postpone those questions. It is designed around them.


Pilot programs need a decision before they begin


A pilot should start with a decision statement, not a project description. The difference is small in wording and significant in practice. “Trial a new customer intake process in two locations” describes activity. “Decide whether the new intake process should replace the current model across all locations” identifies the choice the evidence must support. It gives the pilot an endpoint and tells everyone why the work exists.


The decision statement should name the available outcomes before delivery begins. Scale, adapt and retest, or stop are usually enough. Without these options, teams can unconsciously treat continuation as the default and search for evidence that protects the work already invested. Pre-agreed choices make it easier to interpret mixed results honestly, including the possibility that a sensible idea did not perform well enough under real conditions.


Experimentation is not the same as uncertainty without limits


Experimentation is valuable because it turns uncertainty into questions that can be tested. It is not permission to remain indefinitely unsure. A sound experiment begins with an explicit hypothesis about what will change, for whom and under what conditions. It also acknowledges the assumptions that could make the result misleading, such as unusually motivated staff, extra implementation support or a participant group that does not reflect the wider population.


The OECD’s guidance on experimentation emphasises choosing a method that fits the policy or service question, including the trade-offs around generalisability, resources and causal inference. Organisations do not need to turn every workplace or service improvement into a laboratory study. They do need enough discipline to distinguish evidence from impression. The method should be proportionate to the decision, but still capable of answering it.


Program evaluation should be designed with delivery


Program evaluation is often treated as something that happens after a pilot, when the delivery team hands its results to someone else for assessment. By then, the strongest opportunities to create credible evidence may have passed. Baseline measures may be missing, data definitions may have changed and participants may have been selected for convenience rather than relevance. Evaluation needs to be designed at the same time as the pilot, not attached when the presentation is due.


The Australian Centre for Evaluation advises organisations to establish evaluation objectives, scope, evidence sources and reporting needs early. This means deciding what data will be collected, who owns it, how privacy and cultural considerations will be managed and what limitations will be accepted. It also means capturing qualitative experience. A result can look efficient in a dashboard while creating confusion, exclusion or extra unpaid effort for the people expected to use it.


A small test still needs representative conditions


Pilots often receive advantages that disappear at scale. The project team may provide immediate support, senior leaders may remove obstacles and participants may volunteer because they are already interested in the change. These conditions help a new idea get started, but they can also disguise the effort required to make it work normally. If the pilot depends on exceptional attention, that dependency is part of the result.


Representative does not mean reproducing every possible condition. It means deliberately including enough variation to expose the risks that matter. A workplace pilot should consider different roles, locations, working patterns and levels of confidence. A customer pilot should include people with different needs, access requirements and digital capability. A process pilot should encounter ordinary workloads, handovers and exceptions. Testing only the easiest path tells you that the easiest path is easy.


Pilot success criteria must include scale


Most success criteria describe what happens inside the pilot: uptake, satisfaction, time saved or errors reduced. Those measures matter, but they do not answer whether the model is ready to expand. Scale introduces new questions about training, ownership, technology, procurement, accessibility, support capacity and the cost of maintaining quality. A pilot can deliver a strong local outcome while revealing that the broader operating model is not yet viable.


The evaluation should therefore include scaling conditions as well as pilot performance. What capability must exist in each location? Which benefits depend on the project team remaining closely involved? What cost emerges when temporary workarounds become permanent processes? What risks increase when participation is no longer voluntary? These are not reasons to avoid scaling. They are the evidence needed to scale with fewer surprises and a more credible investment case.


Governance should protect learning, not the pilot


Pilots become difficult to stop when governance is focused on delivery milestones rather than learning. Steering groups ask whether the launch occurred, whether the budget is on track and whether stakeholders are satisfied. Those are reasonable controls, but they can make the pilot itself the thing being protected. Governance should also ask whether the evidence is strong enough, which assumptions have failed and whether continuing the test will materially improve the decision.


This requires psychological and practical permission to report an unfavourable result. If stopping a pilot is treated as failure, teams will naturally keep adjusting the story until continuation appears sensible. A well-run pilot that disproves an assumption can save far more than a weak pilot that produces a positive case for a costly rollout. The governance test is whether leaders value a clear answer more than a reassuring answer.


What a decision-ready pilot includes


The documentation does not need to be elaborate. It needs to make the logic visible and the decision difficult to avoid. A concise pilot brief should include:


  • Decision statement: Name the choice that will be made when the pilot ends and who is authorised to make it.

  • Explicit hypothesis: Describe the expected change, the people affected and the assumptions that must hold.

  • Representative conditions: Select participants, locations and operating situations that expose meaningful variation and risk.

  • Evidence plan: Set baseline measures, data sources, qualitative methods, responsibilities and ethical safeguards before launch.

  • Thresholds and trade-offs: Agree what results would justify scaling, adaptation or stopping, including cost and unintended effects.

  • Owner and decision date: Give one accountable person responsibility for convening the decision by a fixed date.


These elements prevent the pilot from being judged solely by the enthusiasm of the people closest to it. They also improve the conversation when results are mixed. The organisation can see which assumptions were supported, which conditions need work and whether another test would resolve a genuine evidence gap or simply delay an uncomfortable choice.


The final report should end with a choice


A final pilot report should be written for the decision-maker, not as a record of every activity completed. It should state the original decision, summarise the evidence, explain important limitations and recommend one of the pre-agreed outcomes. If adaptation and retesting is recommended, the report should specify what remains uncertain and how the next test will answer it. “More time is needed” is not enough.


Evaluation findings create value only when they influence what happens next. The Australian Centre for Evaluation notes that findings should support decisions about whether a program is changed, continued, replicated or terminated, and that implementation of agreed improvements should be planned. That closes the loop between learning and action. A pilot is not successful because it reached the end of its schedule. It is successful because the organisation is better able to choose.


The practical question for leaders is simple: if this pilot finishes on Friday, what decision will be made on Monday? If nobody can answer, the problem is not yet the quality of the evidence. It is the design of the pilot.


References

bottom of page