Prioritizing CRO Tests: PIE, ICE, and Beyond
On this page
Once your CRO program is running, you quickly hit a pleasant problem: you have more ideas and hypotheses than you have capacity to test. Research surfaces many barriers, the team has many ideas, and you can only run so many tests at once (constrained by traffic, development time, and attention). So you have to prioritise — decide which tests to run first — and how you prioritise largely determines your CRO program’s productivity: prioritise well (test the highest-impact, most promising ideas first) and you get more conversion improvement faster; prioritise poorly (test low-impact or low-confidence ideas, or just whatever’s top of mind) and you waste your limited testing capacity. Prioritisation frameworks like PIE and ICE exist to make this systematic — scoring ideas on consistent criteria so you can rank them objectively rather than by gut or politics. This piece covers how to prioritise CRO tests using PIE, ICE, and sensible judgment beyond the frameworks. (This connects to the CRO guide and hypothesis discussions; this focuses on prioritisation.)
This piece covers why prioritisation matters, the PIE framework, the ICE framework, how to use them well, and the judgment beyond the frameworks. Because you’ll always have more ideas than capacity, and prioritising well is what makes your CRO program productive. Let me walk through it.
Why prioritisation matters
Prioritisation matters because your testing capacity is limited and the ideas aren’t. More ideas than capacity — a running CRO program generates many ideas and hypotheses (from research, the team, observations), but you can only test so many at once (traffic limits how many concurrent tests you can run validly, and development and attention are finite), so you must choose. Capacity is precious — each test takes traffic, time, and effort, and you can only run so many, so each test “slot” is precious — spending it on a low-value test means not testing a high-value one. Impact varies hugely — CRO ideas vary hugely in potential impact (some could meaningfully lift conversion, others are minor tweaks), so testing high-impact ideas first yields far more improvement than testing low-impact ones. Confidence varies — ideas also vary in how likely they are to work (some are well-grounded in strong evidence, others are speculative), so testing higher-confidence ideas first yields more wins. Effort varies — ideas vary in implementation effort (some are easy, some require significant development), affecting how many you can do and how quickly. And good prioritisation maximises output — prioritising by impact, confidence, and effort maximises the conversion improvement you get from your limited testing capacity (more high-value tests, faster). So prioritisation matters because you have more ideas than capacity, each test slot is precious, and ideas vary hugely in impact, confidence, and effort — so prioritising well (testing the highest-value ideas first) maximises your CRO program’s output, while prioritising poorly wastes precious capacity. So take prioritisation seriously — it’s a key determinant of CRO productivity. The frameworks (PIE, ICE) help do it systematically, covered next.
The PIE framework
PIE is a common CRO prioritisation framework scoring ideas on three criteria. P — Potential — the potential improvement: how much could this test improve conversion (how big is the opportunity / how much room for improvement on the page or element)? Higher potential = higher priority. I — Importance — the importance: how important/valuable is the page or element (how much traffic does it get, how valuable is it)? A test on a high-traffic, high-value page (the product page, checkout) is more important than one on a low-traffic page, because the same improvement affects more visitors/value. E — Ease — the ease of implementing the test: how easy is it to build and run (effort, complexity, technical difficulty)? Easier = higher priority (you can do it faster, for less effort). How it works — score each idea on P, I, and E (e.g., 1-10 each), then combine (average or sum) into a priority score, and rank ideas by score — prioritising high-potential, high-importance, easy tests. The logic — PIE prioritises ideas with high potential improvement, on important (high-traffic/value) pages, that are easy to implement — sensibly focusing capacity on impactful, high-value, achievable tests. So PIE scores ideas on Potential (improvement opportunity), Importance (page traffic/value), and Ease (implementation effort), combining into a priority score to rank ideas — prioritising high-potential, high-importance, easy tests. It’s a sensible, simple framework for prioritising by the factors that matter (impact via potential and importance, and effort via ease). So PIE is a useful prioritisation tool, scoring the key factors. The next section covers ICE, a similar alternative.
The ICE framework
ICE is another common prioritisation framework, similar to PIE, scoring on three criteria. I — Impact — the impact: how much could this improve the metric (similar to PIE’s potential — the size of the expected effect)? Higher impact = higher priority. C — Confidence — the confidence: how confident are you that it will work (how strong is the evidence/reasoning behind it — a well-grounded hypothesis has higher confidence than a speculative one, as the hypothesis discussion covers)? Higher confidence = higher priority. E — Ease — the ease of implementation (same as PIE’s ease — effort/complexity)? Easier = higher priority. How it works — score each idea on I, C, and E (e.g., 1-10), combine into a score, and rank — prioritising high-impact, high-confidence, easy ideas. The key difference from PIE — ICE includes Confidence (how likely the idea is to work, based on evidence), which PIE doesn’t explicitly (PIE’s “importance” is about the page’s value, not the idea’s likelihood of working). So ICE emphasises confidence (evidence-grounding), which is valuable — prioritising ideas you’re confident in (well-grounded) over speculative ones. The logic — ICE prioritises high-impact, high-confidence (well-grounded), easy ideas — sensibly focusing on impactful ideas likely to work that are achievable. So ICE scores ideas on Impact (expected effect size), Confidence (likelihood of working, based on evidence), and Ease (implementation effort), ranking by the combined score — prioritising high-impact, high-confidence, easy ideas. Its emphasis on confidence (evidence-grounding) is a useful complement to impact and ease. So ICE is another useful framework, with confidence as its distinctive, valuable criterion. PIE and ICE are similar (both score impact and ease); the main difference is PIE’s importance (page value) versus ICE’s confidence (idea likelihood). The next section covers using them well.
How to use the frameworks well
PIE, ICE, and similar frameworks are useful tools, but using them well requires some judgment. Use them as a guide, not gospel — the frameworks help systematise prioritisation (scoring and ranking ideas consistently), but the scores are estimates (your judgment of potential, impact, confidence, ease), so use the ranking as a guide, not an exact truth — it surfaces the strong candidates, but apply judgment. Score consistently — score ideas consistently (the same way across ideas), so the rankings are comparable; inconsistent scoring undermines the comparison. Combine impact, confidence/value, and ease — whether PIE or ICE (or a custom variant), the key is scoring the factors that matter: impact/potential (how much it could improve things), confidence and/or importance (how likely to work, how valuable the page), and ease (effort) — these are the right factors, so a framework covering them prioritises sensibly. Favour high-confidence and high-impact — prioritise ideas that are both reasonably high-impact and reasonably high-confidence (well-grounded in evidence, as the hypothesis discussion covers), since these are most likely to yield meaningful wins — avoid testing speculative low-confidence ideas just because they’re easy, or chasing high-potential ideas with no evidence. Consider quick wins and big bets — balance easy quick wins (high ease, decent impact — fast value) with bigger bets (high impact, maybe more effort — potentially large value), rather than only doing easy tests or only big ones. And re-prioritise as you learn — re-prioritise as you learn (test results, new research update your estimates of impact and confidence), keeping the prioritisation current. So use the frameworks well by treating them as a guide (not gospel), scoring consistently, ensuring you cover the key factors (impact, confidence/value, ease), favouring high-confidence high-impact ideas, balancing quick wins and big bets, and re-prioritising as you learn. The frameworks systematise prioritisation usefully, but applied with judgment (not mechanically). So use PIE or ICE (or a sensible variant) to systematically rank your ideas, applying judgment to the results, and you prioritise your testing capacity well.
The judgment beyond the frameworks
Beyond the frameworks, some judgment factors matter in prioritisation. Strategic alignment — prioritise tests aligned with business goals and strategy (e.g., if AOV is a priority, prioritise AOV-related tests), not just the highest framework score, since CRO should serve strategy. Evidence strength — weight evidence strength heavily (ICE’s confidence): ideas grounded in strong evidence (clear research showing a barrier) deserve priority over speculative ones, since they’re more likely to win and teach you something real (as the hypothesis discussion covers). Learning value — sometimes prioritise tests for their learning value (a test that, win or lose, teaches you something important about your customers), not just immediate impact, since learning compounds. Dependencies and sequencing — consider dependencies (some tests inform others) and sequence sensibly. Risk — consider risk (a test that could hurt the experience or has downside) and weigh it. Capacity realities — prioritise within your real capacity (traffic for valid tests, development resources), being realistic about how many tests you can run well. And don’t over-engineer prioritisation — the frameworks and judgment help, but don’t over-engineer prioritisation itself (spending more time scoring than testing); the goal is sensibly focusing capacity, not perfect ranking. So beyond the frameworks, apply judgment on strategic alignment, evidence strength (favouring well-grounded ideas), learning value, dependencies/sequencing, risk, capacity realities, and not over-engineering the prioritisation. The frameworks plus this judgment give sensible prioritisation — focusing your limited testing capacity on the highest-value, most promising, strategically-aligned tests. So combine the frameworks (for systematic scoring) with judgment (for the factors frameworks don’t fully capture), and you prioritise your CRO testing well — maximising the conversion improvement from your limited capacity. The point is sensible focus, not mechanical ranking, so use frameworks and judgment together toward that end.
A worked example: ranking a backlog with ICE
Picture a store with a backlog of six CRO ideas and capacity to run only a couple of tests at a time. Scoring them with ICE (Impact, Confidence, Ease, each 1-10) makes the right order obvious. Idea one: add a clear size guide to product pages — strong session-recording and survey evidence that sizing confusion causes abandonment (high Confidence, say 8), a plausibly meaningful effect on add-to-cart (high Impact, 7), and fairly easy to build (high Ease, 8). That’s a strong all-round score. Idea two: redesign the entire homepage — potentially high Impact (7) but low Confidence (3, it’s a big speculative change with weak evidence) and low Ease (3, lots of work). High potential, but risky and expensive. Idea three: change a button colour — easy (Ease 9) but low Impact (2) and low Confidence (3, no evidence it matters). Easy but pointless. Idea four: simplify the checkout’s address step — solid evidence of friction there (Confidence 7), high Impact because checkout is where intent converts (8), moderate Ease (5). Another strong candidate.
Ranking by combined score, ideas one and four rise to the top — both reasonably high-impact, well-grounded in evidence, and achievable — so you test those first. The homepage redesign, despite its high potential, drops down because its low confidence and high effort make it a poor early bet (you’d want more evidence first, and you might break it into smaller testable pieces). The button-colour tweak sinks because, easy as it is, there’s no reason to expect it matters. Notice how the framework prevented two classic mistakes: chasing the exciting-but-speculative big redesign first, and filling capacity with easy-but-trivial tweaks. It directed limited capacity toward impactful, well-evidenced, doable tests. That’s the framework doing its job — and notice the human judgment layered on top (breaking the homepage into smaller testable pieces, recognising the button test as a waste despite its ease), which is exactly how frameworks and judgment work together.
Avoiding the prioritisation traps
A few traps catch teams even when they use a framework, and they’re worth naming because avoiding them is most of getting prioritisation right. The first is the easy-tweak trap: because Ease is one of the criteria, it’s tempting to fill your testing capacity with trivial, easy changes (button colours, minor copy tweaks) that score well on Ease but poorly on Impact, producing a busy testing calendar that moves nothing. Guard against it by ensuring impact and confidence carry real weight, not just ease. The second is the shiny-big-bet trap: falling in love with a dramatic high-potential idea and testing it despite weak evidence and high effort — which often wastes a lot of capacity on something that may not work. The fix is to insist on confidence (evidence) before committing to expensive tests, and to break big ideas into smaller, evidence-gathering steps.
The third trap is treating the scores as objective truth, when they’re really structured estimates — so a team can spend hours debating whether something is a 6 or a 7, or trust a ranking that’s only as good as the guesses behind it. The remedy is to use the scores to surface strong candidates and then apply judgment, rather than following the numbers mechanically. The fourth is over-engineering prioritisation itself: building elaborate scoring spreadsheets and arguing over decimals while running few actual tests. Remember the goal is sensible focus, not a perfect ranking — if scoring is eating time that should go to testing, simplify it. And the fifth is never re-prioritising: setting an order once and grinding through it even as test results and new research change what you should believe about impact and confidence. Re-prioritise regularly as you learn. Avoid these five traps — easy-tweak bias, shiny big bets without evidence, false precision, over-engineering, and stale priorities — and your prioritisation, framework plus judgment, will reliably point your limited testing capacity at the work most likely to move conversion. That, ultimately, is all prioritisation is for: making sure the next test you run is the most valuable one you could run.
The bottom line
Once your CRO program is running, you’ll always have more ideas and hypotheses than capacity to test (traffic, development, and attention are limited), so you must prioritise — and how you prioritise largely determines your program’s productivity: prioritise well (test the highest-impact, most promising ideas first) and you get more conversion improvement faster; prioritise poorly and you waste precious testing capacity. Prioritisation frameworks make this systematic. PIE scores ideas on Potential (how much improvement is possible), Importance (how valuable the page is, by traffic and value), and Ease (implementation effort), combining into a score to rank ideas — prioritising high-potential, high-importance, easy tests. ICE scores on Impact (expected effect size), Confidence (how likely it is to work, based on evidence), and Ease — prioritising high-impact, high-confidence, easy ideas, with Confidence (evidence-grounding) as its distinctive, valuable criterion (versus PIE’s page-value-focused Importance). Both are useful; both cover impact and ease, differing mainly in importance versus confidence. Use them well by treating them as a guide rather than gospel (the scores are estimates — apply judgment), scoring consistently, ensuring you cover the key factors (impact, confidence and/or value, ease), favouring ideas that are both high-impact and high-confidence (well-grounded in evidence, not speculative), balancing easy quick wins with bigger bets, and re-prioritising as test results and research update your estimates. And apply judgment beyond the frameworks: strategic alignment (prioritise tests serving business goals), evidence strength (weight well-grounded ideas heavily), learning value (some tests are worth running for what they teach), dependencies and sequencing, risk, capacity realities (be realistic about how many tests you can run well), and not over-engineering the prioritisation itself (the goal is sensible focus, not perfect ranking, so don’t spend more time scoring than testing). Combine the frameworks (systematic scoring) with this judgment (the factors they don’t fully capture), and you focus your limited testing capacity on the highest-value, most promising, strategically-aligned tests — maximising the conversion improvement your CRO program delivers. Prioritisation is a key determinant of CRO productivity, so do it thoughtfully with frameworks and judgment together.
Frequently asked questions
What is the PIE framework for CRO prioritisation?
PIE is a framework that scores CRO test ideas on three criteria: Potential (how much improvement is possible — how big the opportunity is on that page or element), Importance (how valuable the page or element is, based on its traffic and value — a test on a high-traffic product page or checkout is more important than one on a low-traffic page, because the improvement affects more visitors), and Ease (how easy the test is to implement — effort and complexity). You score each idea on the three criteria (often 1-10 each), combine them into a priority score, and rank your ideas — prioritising high-potential, high-importance, easy tests. PIE sensibly focuses your testing capacity on impactful tests on valuable pages that are achievable, making prioritisation systematic rather than gut-driven.
What is the difference between PIE and ICE?
Both are similar prioritisation frameworks that score ideas on three criteria and rank by the combined score, and both include Ease (implementation effort) and a measure of impact (PIE’s Potential, ICE’s Impact). The key difference is the third criterion: PIE uses Importance (how valuable the page is, by traffic and value), while ICE uses Confidence (how likely the idea is to work, based on the strength of the evidence and reasoning behind it). ICE’s emphasis on Confidence is valuable because it prioritises well-grounded ideas (backed by strong evidence) over speculative ones, which tend to yield more wins. In practice, the best prioritisation often considers all of these factors — impact, page value/importance, confidence, and ease — so you can use either framework, or a sensible variant covering all of them.
How should I actually use these frameworks?
As a systematic guide, applied with judgment — not as gospel. Score your ideas consistently on the criteria (so the rankings are comparable), combine into scores, and use the ranking to surface your strongest candidates. But remember the scores are estimates (your judgment of potential, impact, confidence, and ease), so apply judgment to the results rather than following them mechanically. Favour ideas that are both reasonably high-impact and reasonably high-confidence (well-grounded in evidence), balance easy quick wins with bigger higher-impact bets, and re-prioritise as test results and new research update your estimates. Also weigh factors the frameworks don’t fully capture — strategic alignment with business goals, learning value, dependencies, and risk. The goal is sensibly focusing your limited testing capacity, not producing a perfect ranking, so don’t over-engineer the scoring itself.
What factors matter most in prioritising CRO tests?
The core factors are impact (how much the test could improve conversion or the relevant metric), confidence (how likely it is to work, based on the strength of the evidence behind it), and ease (how much effort it takes to implement) — with page importance (traffic and value) also mattering, since the same improvement is worth more on a high-traffic, high-value page. Beyond these, weigh strategic alignment (prioritise tests serving your business goals), evidence strength (well-grounded ideas deserve priority over speculative ones, as they’re more likely to win and teach you something real), and learning value (some tests are worth running for what they’ll teach you about your customers). Balance quick wins against bigger bets, be realistic about your true testing capacity, and re-prioritise as you learn — focusing capacity on the highest-value, most promising, strategically-aligned tests.
