STR.7 - Decide Whether and How to Experiment While Preserving Strategic Options
Type: Method pattern Status: —
STR.7:1 - Problem frame
Use this pattern when a consequential strategic question remains and someone proposes a trial, pilot or experiment to answer it. First ask which attainable result could change the choice or warranted claim, and whether obtaining that result is worth its full burden.
The first useful result may be to use existing evidence, narrow a proposal, retain a possibility without testing it or stop. When an experiment is worthwhile, the result is a bounded design that connects possible observations to a decision, with the needed authority, resources, protections and stopping conditions. Designing it does not establish that it has been authorized or performed.
The practical gain is to buy useful evidence without making “run a pilot” the compulsory response to uncertainty. The scope is strategic use of an experiment; qualified domain methods retain experimental design, measurement, causal inference and any necessary professional assurance.
Do not add an experiment when a sufficient current comparison already answers the question. A required permission cannot be acquired by exposing people or data first and calling the exposure a test. A retained option may deserve memory or maintenance without a new investigation.
STR.7:2 - Problem
A pilot can become a small implementation of the favoured strategy rather than a discriminating inquiry. Positive anecdotes justify expansion; adverse results trigger another pilot; no result can end the proposal.
The apparent price also hides preparation, participant time, support, interpretation, restoration and displaced work. A “small” test may consume the only people available for a protected service. Conversely, refusing every experiment because evidence is imperfect can miss an affordable answer that would prevent a much larger mistake.
STR.7:3 - Forces
Further evidence can improve choice, but delay and exposure can make the attainable investigation worse than acting on a sufficient present basis. A narrow experiment is easier to bound, while its result supports a narrower claim.
Preserving several possibilities can protect future choice but incurs carrying cost. Testing one can improve its evidence while closing others through scarce resources or a new obligation. Inquiry and option retention therefore need related but separate comparisons.
STR.7:4 - Solution
Compare inquiry with the feasible alternatives before designing a trial. For a worthwhile experiment, make the decision link, full burden and protected boundary operational. Return the design, a sufficient non-experiment answer or the obstacle.
STR.7:4.1 - State the decision and the claim the experiment could change
Recover the strategic subject, horizon and live alternatives. Write the relied-on claim at the scope actually needed: for example, whether a specified service can be delivered under stated customer terms, not whether “the market likes service”.
Describe materially different possible answers and what each would change. An answer can reject the option, narrow its scope, support a further bounded commitment or leave a decisive gap unresolved. A study can also earn its cost by establishing a warranted explanation or opening a relevant option at a longer research horizon. State that gain rather than inventing an immediate operational decision.
Use a sufficient PSD.10 uncertainty account directly. If every plausible result leaves the present choice and warranted use unchanged, identify another useful question or stop. Statistical precision alone is not the strategic contribution.
STR.7:4.2 - Compare the attainable inquiry with cheaper continuations
Check existing records, calculations and qualified direct results first. Compare the proposed experiment with using that basis, asking a narrower question, making a reversible bounded choice, deferring or stopping. FPF C.11 governs this choice; C.11.DUA helps expose the work hidden in a demand for more evidence.
Include the full increment to the present configuration: design, execution, participation, checking, interpretation, support, recovery, delay and displaced work where material. Keep costs imposed on customers or other participants visible even when they fall outside the sponsor’s budget.
Use qualitative comparison when it is enough. A numerical value-of-information model needs defensible inputs; uncertain probabilities need not be invented to obtain a decision. If a plausible burden or benefit difference reverses the recommendation, examine that difference before commissioning the inquiry.
A binding evidence or protection condition remains in force. Its merits may warrant reconsideration by the responsible authority, but an attractive experiment does not remove it.
STR.7:4.3 - Preserve the alternative being displaced
Name the serious rival to experimenting now. It may be continued operation, a smaller investigation or another use of the same people. State what is forgone, delayed or made harder to recover.
Keep retention separate from testing. Preserving a short account of an unchosen direction is different from paying to retain its capability, contract or interface. STR.6 supplies the retention reason and exercise conditions; STR.9 can compare the resulting flexibility. A stepping stone can remain untested until its next contribution warrants work.
If the question requires an explicit policy for a still-live collection of exploration lines, use C.19 to decide which remain worth exploring, which warrant retention without new tests and which no longer warrant a place in that collection. State the chosen policy, its intended contribution and comparison basis, and the change that would justify reconsideration. This answers the continuing search-policy question; sections 4.1–4.2 compare the worth of one proposed experiment.
STR.7:4.4 - Design the smallest experiment that can supply the useful answer
Start from the needed observation and work backwards to what must actually be offered, used or performed. Choose a form that preserves those conditions with the least full burden. A minimum viable test is the smallest sufficient way to obtain that answer, not necessarily a smaller version of the final product. These examples distinguish different questions, not successive maturity levels:
| Question the result must answer | A possible test construction | What its result can support |
|---|---|---|
| Will these people take the offered next step? | Show a comprehensible offer or demonstration and provide an observable opportunity to accept or decline that step. | A response to that offer under those conditions. Viewing or joining a waiting list alone does not establish paid or continued use. |
| Will they use or pay for the proposed service under these terms? | Provide a limited real service, using competent manual work where that is permitted, and observe use, payment and delivery effort. | The service with that actual support. Human delivery does not establish the cost or performance of a future automated arrangement. |
| Can the proposed technical mechanism perform the required work? | Exercise that mechanism under the relevant conditions, or use qualified evidence that answers the same claim. | Performance within the tested or qualified conditions. A human substituting for the mechanism cannot establish that mechanism’s performance. |
Remove a feature or substitute a simpler component only while the required observation and its interpretation remain possible. Retain the work needed to collect the observation, protect participants and close the activity. Explain what participants will receive and who may handle their information; do not sell simulated performance as a delivered capability.
Use the appropriate domain design to connect observations to the bounded claim. State the participants or systems exposed, contribution being tried, conditions of use, observations, interpretation and limits. Specify comparison or control conditions when the inference requires them; merely naming a pilot does not establish a causal effect.
When the decision specifically needs the effect of changing an offer or feature, one possible design is a concurrent randomized comparison. Construct it as follows, with competent domain and statistical support where the inference needs it:
- State the change and comparator, the people or systems to which the answer must apply, the outcome and any material harms. Choose a meaningful difference and observation period from the receiving decision. Keep the other conditions compatible with the intended comparison.
- Define the units to be assigned, then allocate eligible units to the variants by a random procedure under stated probabilities before exposure. Keep a unit’s assignment consistent unless the design explicitly handles switching. Repeated visits by one person are not automatically independent units; interaction between participants or competition for shared resources can require a different design. Sending each offer to half the recipients does not establish random assignment.
- Specify the sample size needed for decision-relevant precision, whether it is attainable, the outcome calculation, missing-observation treatment and analysis or stopping rule before inspecting the effects. Retain assignment, actual exposure and outcome evidence for both variants, including departures from the plan. Check whether observed group sizes and missingness are compatible with the allocation and collection design; resolve a material unexplained discrepancy before relying on the effect. Judge the resulting difference with its uncertainty and supported population, configuration and period.
This construction starts a controlled comparison; it does not make every pilot an A/B test or supply a qualified causal estimate by itself. C.28 governs that stronger use. Use another qualified design or an adequate existing result when it answers the question better, and return an unattainable inference rather than exposing participants to an uninformative test.
Before dependent activity, identify who may authorize participation, access, spending and changes to existing provision. Set the resource ceiling, duration or stopping event, exception response and withdrawal or restoration conditions. The ceiling includes the work needed to close the experiment responsibly.
Agree what a positive, negative or unresolved result would support. A successful local delivery may support a larger bounded comparison, not a general demand forecast. A failed result may reject this configuration without disproving every possible direction.
Keep missing conditions explicit. If the design cannot be completed until permitted use or support is established, return that bounded gap and preserve independent work. A proposed experiment is not an available action merely because its question is valuable.
STR.7:4.5 - Return the design or sufficient stop
Return the supported comparison and selected next contribution. If no experiment is worthwhile, finish with the usable answer and its material limits.
If a design is proposed, the authorized chooser can approve its bounded scope or return it. Actual execution belongs to the capable participants and applicable professional methods. After execution, compare the obtained observations with the claim and decision specified, retaining deviations and limits. STR.12 supports reconsideration of the affected commitment; completing the inquiry does not cancel continuing service or recovery duties.
Reopen when the decision, attainable evidence, affected exposure, authority, burden or option value changes. Do not continue an inquiry merely because it has already consumed effort.
STR.7:5 - Archetypal Grounding
STR.7:5.1 - SensorCo can obtain a sufficient negative answer without a trial
In this constructed case, SensorCo’s existing contracts can support current service for twelve months. Two of six interviewed customers may discuss a bounded paid service trial next quarter at a price ceiling, subject to agreement and permitted use. Six expressions of interest are not a market conversion estimate.
Before the Board’s next trial decision, four already-funded internal days can complete delivery costing from existing records and specify a possible trial. The work needs no new purchase or new customer-data use. The alternative use of those days is a device-diagnostic improvement.
The Board’s supplied criterion protects current service and reserve while seeking a repeat paid contribution less exposed to generic-inspection price competition. Under the stated conditions, costing can change whether the service is worth pursuing. The Board is willing to postpone the diagnostic work for that attainable answer. This justifies preparation, not yet trial execution or a twelve-month service direction.
The result branches are concrete:
| Attainable preparation result | Supported strategic return |
|---|---|
| Complete travel and incident-support costing shows no relevant permitted service configuration can cover cost at attainable customer terms | Recommend viable device-only continuation for this horizon without a trial |
| A feasible bounded configuration remains and an attainable trial could change the next commitment | Return the bounded trial design, full burden and missing authorizations for separate choice |
| A decisive cost component or required permitted-use answer is missing | Return that gap; do not treat the unknown as zero cost or permission |
A trial could later test the two consenting customers’ use and payment under their agreed terms and the specified delivery arrangement. It cannot establish general market demand or substitute for a required data-use determination. The four preparation days also do not approve the later eighteen-person-day development configuration.
STR.7:5.2 - A small experiment design with a bounded claim
Consider a separate constructed professional-practice case. A practitioner already has the capability to provide a diagnostic consultation and eight discretionary hours this week after current obligations. Two prospective clients have agreed to consider one paid session each. Their ability to enter that agreement, permitted inputs and the practitioner’s scope are supplied conditions of this example.
The live decision is whether to offer the same bounded format for the following four weeks. The practitioner proposes at most four hours this week: two hours per client including preparation, the session, recording the result and agreed follow-up. Both sessions finish within seven days. Each two-hour allowance reserves its last fifteen minutes for returning the result or an explicit limitation and closing the agreed activity. Stop exploratory work while that closure time remains. No additional spending, continuing support or consequential change to either client’s operations is promised in this test.
Before starting, state the bounded claim: each participating client will choose to pay for this format at the stated offered price, and the practitioner can deliver the agreed diagnostic result within the two-hour allowance. Observe acceptance or refusal, actual work time and whether the promised result was delivered. A refusal is useful evidence, not a reason to replace the participant silently.
If both deliveries meet these conditions, the result supports considering the same limited offer next month. If delivery exceeds the allowance, revise the format or decline continuation; willingness to pay does not fix capacity. If inputs prevent delivery, return that limitation rather than reporting a successful test. Two sessions support only these bounded claims, not a population conversion rate.
The rival is using those four hours to improve diagnostic guidance for existing clients. This test is sensible only if the bounded answer can improve the coming offer enough to justify that sacrifice. It is not a universal recommendation to every professional.
STR.7:5.3 - A demonstration, a manual service or a working mechanism?
In this constructed case, a small research service is considering a daily briefing of public reports. The immediate decision is whether to offer the same manually supported format for one further week, not whether to finance automation. Two clients have agreed to consider the pilot under stated prices and terms. Permitted inputs, competent staff and authority for the bounded offer are supplied premises.
A video can show the proposed format and invite a response, but the decision needs evidence of actual delivery and a later order on the stated paid terms. The team therefore proposes five daily briefings for each client, produced manually, followed by a real opportunity to order the same format for the next week. It observes delivery against the agreed content, staff effort and acceptance or refusal of that next offer. If no next offer was made, the repeat-order question remains unanswered.
The proposal allows ten staff hours over seven days: two for preparation, five for delivery, one for interpretation and two for participant follow-up and closure. Stop taking on further activity while the remaining allowance still covers closure. The rival is ten hours of improving existing clients’ guidance; proceed only if the attainable answer warrants that sacrifice. Neither unused time nor permission alone establishes that worth.
If both clients receive the agreed briefings within the allowance and order again on the stated terms, the result supports considering this small manual offer with its measured burden. It does not establish population demand or automated delivery. If the actual decision instead requires briefings produced without staff research, a manual service cannot answer it. Use a qualified result about the proposed mechanism or design a bounded technical test; do not add software merely to make the original customer-use test look more complete.
STR.7:6 - Bias-Annotation
Sponsors can prefer observations that preserve their idea. Specify the adverse and unresolved returns before the experiment and retain refusals, departures and missing observations at their actual meaning.
A low sponsor cost can conceal unpaid participant labour or exposure. Ask who bears the burden and who can decline. Small sample size limits the inference; a polished demonstration does not repair it.
STR.7:7 - Conformance Checklist
- The inquiry has an attainable claim, explanatory gain or option-creating contribution at a named horizon.
- Materially different results change the choice or warranted use.
- Existing evidence and feasible lower-burden continuations have been compared.
- Full burden, displaced work and the serious rival remain visible.
- Retention of a possibility is distinct from testing or selecting it.
- A proposed design states exposure, authority, resource and time limits, interpretation and responsible closure.
- The chosen test form supplies the needed observation; substituted components remain visible in the result’s limits.
- Results apply only within their supported conditions.
- A sufficient current answer or negative result can finish without another experiment.
STR.7:8 - Common Anti-Patterns and How to Avoid Them
The pilot that cannot fail. Every observation leads to expansion or another pilot. Name the result that rejects or narrows the proposal before proceeding.
Preparation presented as commitment. Four days of costing become approval for service delivery. Preserve the separate decision and conditions for execution.
Cheap because others pay. Customers supply extensive data preparation and support. Include their burden and agreement in the comparison.
A test that creates its own missing permission. Data are used before the right is established. Obtain the required determination first; inquiry does not remove that boundary.
STR.7:9 - Consequences
A strategic team can stop weak proposals with sufficient evidence and commission stronger inquiries with a clear purpose. Useful diversity remains available without obligating work on every possibility.
The result may still be inconclusive. That can be a truthful completion when the inquiry reaches its supported limit. Additional work needs a new attainable contribution; sunk effort is not that contribution.
STR.7:10 - Rationale
An experiment is one way to improve a decision’s basis, not a universal stage of Strategy. The relevant comparison concerns attainable evidence and its whole cost, including delay and alternative uses of scarce means.
Design, authorization, performance and inference remain separate because each can fail independently. Keeping them distinct lets a practitioner finish a useful design or negative answer without overstating what has happened.
STR.7:11 - SoTA-Echoing
When is experimental learning worth further commitment? Camuffo and colleagues’ 2024 replication and extension studies a scientific approach in four randomized trials involving 759 firms and reports more idea termination with a non-linear pattern of pivots. Adapt the decision-linked hypothesis and termination contribution in sections 4.1 and 4.5. Reject “always build and test” as the default: sufficient existing evidence can justify stopping. The studies concern the reported training interventions and firms, not a guarantee of revenue or proof that every strategic question needs an experiment.
For accountable experimentation, adapt the OECD’s 2024 policy brief and 2025 discussion of policy experimentation. Their attention to authority, capacity, evaluation and later disposition changes section 4.4’s bounded design and section 4.5’s separate continuation. Policy cases inform this adaptation; they do not replace domain participation or assurance methods.
For constructing a purpose-fit test, adapt the question-first move in Ries’s The Lean Startup (2011, Part Two introduction and chapters 5–6) and his MVP guide, alongside current GOV.UK prototyping guidance. Section 4.4 works backwards from the needed observation and varies what must actually function; section 5.3 makes that choice usable. Reducing a final product’s feature list is a weaker default when the remaining features cannot answer the question. A qualified existing result remains cheaper when it suffices. The startup accounts are historical illustrations, not general effectiveness estimates; the government guidance distinguishes prototype interaction from production readiness in its service-design setting. Neither supplies a technical performance or participation result for this experiment. Reopen when the inference needs different fidelity, exposure or support.
For attributing a change to an intervention, adapt the concurrent comparison in Ries’s chapter 7, with NIST’s randomized-design guidance. Section 4.4 makes allocation an actual random procedure, not a label for successive customer groups. The 2023 review by Larsen and colleagues, especially sections 1.2 and 4–6, qualifies inference through the experimental unit, observation horizon, analysis and interference. Its online-experiment Methods need their stated assumptions; the book’s reported successes do not establish general effectiveness. Reopen the design when those conditions no longer support the required effect claim.
C.11.DUA supplies the full-burden comparison against a sufficient answer or cheaper continuation. This preserves a useful experiment when it can change a consequential choice, while avoiding work whose plausible results would not improve the use. Reopen when evidence, exposure or attainable means change that comparison.
STR.7:12 - Relations
STR.6 supplies option and retention questions; STR.9 and STR.10 compare flexibility and commitment consequences. STR.7 adds the strategic inquiry-worth and bounded-design contribution. STR.11 distinguishes retained alternatives, selected joint use and authorized commitments; STR.12 uses the actual result in reconsideration.
PSD.10 supplies qualified uncertainty when needed. FPF C.11 governs local choice and C.11.DUA the burden of evidence demands. C.19 supplies the continuing exploration-policy result described in section 4.3. The direct experimental and domain methods supply their qualified evidence and performance results.
STR.7:End