A growth experimentation process is a backlog of written hypotheses, scored the same way every time, with success criteria agreed before launch and every result recorded in a learning log. That structure separates experimenting from trying tactics at random, and it works for a team of two running a handful of tests a month. The compounding comes from the log: each test makes the next one better informed.
What counts as a growth experiment, and what does not
A growth experiment is a deliberate change made to answer a question, with four parts written down before it starts:
- A prediction: what you expect to happen and why.
- One primary metric the change should move.
- A comparison: a control group, a holdout or, at minimum, a documented baseline.
- A decision rule: which result means ship, kill or iterate.
If one is missing, the work may still be worth doing, but it isn’t an experiment and shouldn’t count toward your testing pace.
| Looks like an experiment | What it actually is | Where it belongs |
|---|---|---|
| “Let’s try TikTok and see” | A channel launch with no baseline or kill criteria | Backlog, once it has a budget cap, a stop date and a metric |
| Redesigning the homepage | A project with dozens of changes at once | Roadmap; test the riskiest single change separately |
| Fixing a broken form or slow checkout | A fix you would ship regardless | Fix list, shipped this week |
| Copying a competitor’s pricing page | A guess with no evidence of why it works for them | Backlog, only after it has a written hypothesis |
Teams that push fixes through the testing process slow everything down and inflate their win rate. If a change has no plausible downside and you’d ship it whatever the data said, just ship it.
Writing a hypothesis tied to one metric
A hypothesis forces you to state what you saw, what you’ll change and what should move. I use this template:
Because we saw [evidence], we believe [change] for [audience] will increase [primary metric]. It’s a win if [metric] improves by at least [minimum effect] over [duration], without [guardrail metric] getting worse.
A weak version: “A new hero headline will improve conversions.” A usable one, as a hypothetical example: “Because sales call notes show most prospects ask whether we integrate with their CRM, we believe adding integration logos and a one-line answer above the fold for paid search visitors will increase demo request rate. It’s a win if demo request rate rises by at least 15% relative over four weeks, without the share of unqualified demos going up.”
Three rules keep hypotheses useful:
- Evidence first. Sales calls, support tickets, session recordings, funnel drop-offs and customer interviews count as evidence. “I think” does not.
- One primary metric. Pick the metric closest to the change that still matters to the business. A pricing page test should be judged on plan starts, not next quarter’s revenue.
- Guardrails named up front. A lead form test that lifts submissions but floods sales with junk is a loss.
The primary metric should ladder up to the company’s main outcome. If you haven’t defined the input metrics that drive it, start with choosing a North Star metric and its inputs. Every experiment should target one of those inputs.
Prioritizing the backlog: ICE, PIE and consistent scoring
When ideas outnumber capacity, you need a ranking. The two common models:
| ICE | PIE | |
|---|---|---|
| Factors | Impact, Confidence, Ease | Potential, Importance, Ease |
| Best for | Mixed backlogs across channels | Website pages and funnel steps |
| Strength | Confidence rewards ideas backed by evidence | Importance weights pages by traffic and value |
| Weakness | Impact is easy to inflate | Says nothing about how sure you are |
Either works. The model matters far less than scoring consistently. If one person’s 7 is another person’s 4, the ranking is noise and the loudest idea wins anyway.
Write anchor definitions for every score and keep them in the backlog itself. For ICE on a 1-10 scale:
- Impact: 10 = could move the primary metric across the whole funnel; 5 = affects one segment or step; 1 = marginal.
- Confidence: 8-10 = supported by quantitative data and a related past test; 4-7 = qualitative evidence such as call notes or recordings; 1-3 = opinion.
- Ease: 10 = live in under a day with no developer; 5 = about a week of one person’s time; 1 = needs several teams or a month.
Then add three habits. Score new ideas as a group in a short session, not alone. Have one person moderate so the anchors get applied the same way. Re-score the top of the backlog quarterly, because Confidence should change as the learning log fills up.
Agreeing success criteria and test length before launch
Most arguments about test results are really arguments about criteria nobody wrote down. Settle them before launch in a one-page test brief:
- Hypothesis written in the template, with the evidence linked
- Primary metric and exactly where it is measured
- Guardrail metrics and the level that stops the test
- Minimum effect worth detecting, and the sample size it requires
- Test length in full weeks, covering at least one business cycle
- Decision rule for a win, a flat result and a loss
- Owner, and the date results will be reviewed
Test length comes from sample size, not patience. Decide how many visitors or conversions you need, run until you reach it and don’t stop early because the dashboard looks exciting on day three. The math, and what to do when the numbers don’t add up, is in A/B test sample size and significance.
A reasonable default decision rule: a win gets shipped and logged; a flat result keeps whichever version is simpler to maintain; a loss gets reverted with a note on why; a guardrail breach stops the test immediately.
Experiments beyond the website
Landing page tests are the most visible kind, but often not the most valuable. The same process works anywhere you can split an audience or hold a baseline.
| Area | Example hypothesis (hypothetical) | Primary metric | How to split |
|---|---|---|---|
| Ads | A risk-reversal offer beats a feature-led angle for cold audiences | Cost per qualified lead | The ad platform’s built-in experiment or A/B test tool |
| A shorter onboarding sequence with one action per email improves activation | Activation rate | Random split inside the flow, plus a small no-email holdout | |
| Pricing | Showing annual billing by default raises the annual plan share | Share of new plans on annual | Random split of new visitors on the pricing page |
| Sales process | Calling inbound leads within minutes raises meetings held | Meeting held rate | Alternating lead assignment |
Pricing tests touch trust: showing two people different prices for the same plan at the same time can backfire if they compare notes, so many teams test presentation, packaging or trial terms instead, or change the price for new signups and compare cohorts. Sales process tests only work if reps follow the rule, so keep the rule simple and check compliance weekly. In ads, change one variable at a time; a new audience with a new creative tells you nothing about either.
Treating site, ads, lifecycle and sales as one testable system is how I approach growth strategy and go-to-market work. The best lever is often outside the channel the team usually tests.
The learning log: making results reusable
The backlog decides what to test. The learning log is what makes testing compound. Without it, teams re-run ideas that already failed and new hires repeat old mistakes.
Keep one row per completed test with these fields:
- Test ID, name and dates
- Hypothesis, audience and channel
- Result with actual numbers, and whether it reached the planned sample size
- Decision: shipped, killed or iterating
- The learning: one sentence that could apply beyond this test
- Tags: funnel stage, lever (message, offer, friction, social proof, timing) and audience
- Links to the variants and the brief
The learning sentence is the hardest field and the most valuable. “Variant B won” is a result. “Cold visitors respond more to proof that setup is fast than to feature depth” is a learning you can reuse in ads, email and sales scripts.
Review the log once a month. Look for levers that won more than once across channels and levers that keep coming back flat, then feed those patterns into Confidence scores. Expect plenty of flat and losing results. They are normal, and logging them keeps you from paying for the same lesson twice.
Experiment velocity: how many tests a small team can realistically run
Velocity is capped by three things: traffic or volume, build capacity, and the time it takes to analyze and decide. More ideas don’t help when one of those is the bottleneck.
As a planning range, not a benchmark, a team of two with part-time design and development help can usually complete two to four experiments a month across all channels. A team with a dedicated developer and healthy traffic might finish six to ten. On a low-traffic site, one test per page at a time is often the limit, which is another reason to spread tests across ads, email and sales.
A simple rhythm keeps it moving:
- Weekly, 30 minutes: review running tests against their briefs, close any that reached sample size, log the results and give the top backlog item an owner and launch date.
- Monthly: score new ideas and review the learning log for patterns.
Count velocity as tests completed with a decision and a learning logged, not tests launched. Launching is easy. Finishing, deciding and writing down what you learned is the part that compounds. Keep tests on the same page and audience from overlapping, and keep most of the portfolio in small, fast tests with one or two bigger swings per quarter.
Get it built
If your team runs tests but the lessons don’t stick, I can set up the backlog, scoring anchors, test briefs and learning log, then run the first cycle with you. Start with a Growth Audit, $1,500 fixed and credited if we continue. See pricing or get in touch.