An A/B test can look convincing even when the difference is purely random, especially on an SME website with few conversions. This article explains what evidence a web agency should provide and when other methods support more reliable decisions.
Are you paying for an A/B test that can never produce a reliable answer? If so, the test should not begin. In A/B testing, the crucial question is not how many visits the website receives, but whether the eligible traffic can generate enough measurable conversions to detect a change that is genuinely worth acting on.
This becomes particularly clear on an SME website that receives a respectable number of sessions but only a small number of purchases or quote requests. When traffic is divided between the original and a variant, the number of business events in each group falls even further. Temporary fluctuations can then look convincing in the experimentation tool even though the evidence is too weak to support a confident decision. The result may be unnecessary design, development and launch costs for a solution that does not improve sales—and may even make them worse.

Traffic is sufficient only when A/B testing can detect a relevant change
There is no universal minimum number of visitors that makes a website ready for experimentation. Feasibility depends on the baseline conversion rate, the smallest effect the business needs to detect, the selected statistical method and the amount of qualified traffic that can be allocated between the variants. The same traffic volume may therefore be sufficient to test a major change on a high-converting page but wholly inadequate for detecting a small improvement in a rarely used quote flow.
The first thing we examine is the baseline conversion rate: the proportion of relevant visitors who currently complete the test’s primary goal. The baseline must be based on comparable traffic and reliable tracking. If it combines campaign spikes, bot traffic, internal sessions or a period when the form was broken, the sample size calculation will also be misleading.
Consider a clearly hypothetical example in which a test page has a baseline conversion rate of 2.0 percent and the business wants to be able to detect an increase to 2.4 percent. The difference is 0.4 percentage points, but the relative improvement is 20 percent. The web agency should report both figures, because the phrase “20 percent improvement” can otherwise sound much larger than the absolute change in the number of actual transactions.
The minimum detectable effect, commonly abbreviated to MDE, should not be selected solely to make a calculator show a convenient test duration. It should represent the smallest change that justifies the cost and risk of building, quality-assuring and maintaining the solution. If the business would only change the website in response to a clear sales impact, the test must be sized for that specific decision threshold.
Total website traffic is also a poor measure of the data available for an experiment. Visitors who never reach the test page cannot contribute. The same applies to traffic excluded because of bots, internal users, the wrong market or device types that cannot be exposed to the variant. If the test applies only to the mobile version of a React flow or a particular WordPress template, that is the traffic that should be included in the calculation—not the website’s total session count.
The unit of analysis must also be clearly defined. If the same person is counted as several new sessions and switches between variants, the comparison may be contaminated. A robust setup therefore explains whether randomization occurs by user, account or session and how consistent exposure will be maintained. For a business, this provides better control over whether a difference is caused by the design or by the way the tool allocated visitors.
The calculation should show a coherent sequence: together with the significance level and statistical power, the baseline and minimum relevant effect determine the required sample size per variant. The sample size is then divided by the qualified weekly traffic, adjusted for traffic allocation, to estimate the test duration. This also makes it possible to estimate how many primary conversions the test will realistically need to collect.
This is where the commercial dividing line lies. A test may be statistically feasible if it runs for long enough, yet still be a poor project. If the website can only detect, within a reasonable period, an effect considerably larger than the change could realistically be expected to produce, low-traffic A/B testing becomes more of a lottery than a decision-making tool.

Require a calculation plan before building any variant
A serious test brief does not begin with the color of a button. It begins with the decision the experiment should support and the assumptions required to make that decision meaningful. The agency should be able to present the required sample size, expected test duration and stopping rules before time is spent on design and development.
The table below shows what such a brief might look like. All figures belong to the same hypothetical example and should not be interpreted as general recommendations. The actual values must be calculated for the specific website and the statistical method being used.
| Field | Illustrative content | What the web agency should explain |
|---|---|---|
| Primary conversion | Completed quote request | Why this outcome represents the business objective and how the event is quality-assured. |
| Baseline | A hypothetical 2.0 percent of qualified, unique visitors | Which historical period, data source and denominator were used to establish the rate. |
| Minimum detectable effect | A hypothetical increase to 2.4 percent, equivalent to 0.4 percentage points or 20 percent in relative terms | Why this change is large enough to justify a business decision. |
| Significance level | A hypothetical planning choice of 5 percent in a two-sided test | How the choice affects the risk of a false-positive result and the required sample size. |
| Statistical power | A hypothetical planning choice of 80 percent | How the probability of detecting the planned effect influences the size of the test. |
| Variants and allocation | A control and one variant with an even traffic split | That additional variants or an uneven allocation require a new calculation. |
| Eligible weekly traffic | A hypothetical total of 1 000 qualified visitors, or approximately 500 per variant | Which traffic has been excluded and why it cannot contribute to the test. |
| Calculated sample size | In this hypothetical setup, the requirement could be slightly more than twenty thousand visitors per variant | The exact output from the selected calculator and how the statistical method affects the calculation. |
| Estimated test duration | With the example traffic above, the test would need to run for more than forty weeks before adding further margins | Whether the period is commercially reasonable and whether the traffic mix can be expected to remain comparable for that long. |
| Stopping rule | Analysis after reaching the planned sample size and covering complete business cycles, with reassessment following major changes to tracking or the offer | When the test may be completed, discontinued or reported as having no clear result. |
Planning choices such as a 5 percent significance level and 80 percent statistical power are common, but they provide no guarantees. In a traditional frequentist test, the significance level controls the tolerance for false-positive results when there is no real effect and the test rules are followed. It does not mean that the winner shown by the tool automatically has a 95 percent probability of being the best option.
Similarly, statistical power describes the test’s ability to detect the planned effect if that effect genuinely exists. Greater power generally requires more data, while a lower level increases the risk of failing to detect a relevant improvement. The agency should translate this into decision-making consequences: how much uncertainty the business will accept and what an incorrect implementation or missed improvement opportunity could cost.
The primary outcome must be defined before the test goes live. For a quote-driven business, this might be a completed request that has been recorded correctly from a technical perspective. Button clicks, form starts and scrolling may still be tracked as diagnostic metrics, but they must not be allowed to replace the business objective retrospectively simply because one of their graphs happens to look positive.
The number of variants also affects the ability to make a decision. Every additional version divides the traffic further and creates more comparisons in which random differences may appear interesting. A website with limited data therefore rarely benefits from simultaneously testing several headlines, layouts and button labels without a very strong justification.
The experimentation tool may use a traditional fixed-sample test, a sequential method that permits checks during the test or a Bayesian model. These methods address uncertainty differently and require different stopping rules. A green winner indicator in the interface is therefore no substitute for the agency explaining which method is being used, what assumptions it makes and what will happen if the result remains inconclusive.
AI features can formulate hypotheses, generate variants and summarize experiment data more quickly. They cannot create more qualified visitors or conversions. When the available data is limited, faster production may instead increase costs by allowing more underpowered tests to be built before anyone has assessed whether they can be evaluated.
Test duration must not become a way to wait for a winner
A test period must be long enough to capture normal variation, but short enough for the comparison to continue reflecting the same business environment. Days of the week, pay cycles, sales cycles and the delay between the first visit and conversion can all affect the outcome. It is therefore not enough simply to divide the sample size by an average daily volume and select the earliest possible end date.
The opposite pitfall is allowing the test to continue until statistical significance appears. Over a very long period, campaigns may begin, prices may change, product ranges may be replaced and the traffic mix may shift. An update to WordPress, the React application, consent management or tracking may also change what is recorded without the variant itself becoming better or worse.
Repeated checking is another problem. If the team looks for a positive spike every day and stops the test as soon as one appears, it is no longer using the same decision rule on which the original calculation was based. Sequential methods can support ongoing analysis, but that specific method and its boundaries must be selected from the outset.
This is how we would approach the timing question: document which complete business cycles the test needs to cover, estimate when the planned sample size can be reached and list the changes that would make the comparison difficult to interpret. If the schedule extends across periods in which the offer and traffic are likely to change significantly, the project should not be forced through. Reporting “no clear result” is more professional than manufacturing a winner.
When conversion volume is low, funnel analysis and usability testing often support a better next decision
Insufficient data for an experiment does not mean decision-making must stop. When final conversions are scarce, a funnel analysis can show where users drop off. Moderated usability testing can then explain why they get stuck. These methods do not prove that a particular change will produce a specific conversion increase, but they can provide a much stronger basis for prioritizing sources of friction.
A useful funnel breaks the journey down into observable steps. For a quote flow, these might include the landing page, service page, form start, validation error displayed and completed request. If the analysis only compares all visits with final conversions, it cannot distinguish between a weak landing page, an unclear offer and a technical problem in the form.
GA4 or Matomo can be used to map drop-offs, while Microsoft Clarity can provide session recordings and interaction patterns showing what users did immediately before stopping. However, these tools cannot reliably reveal what the person thought, lacked or misunderstood. A quick return may be caused by anything from the wrong target audience to an essential piece of information being impossible to find.
Segmenting by device type, traffic source and new or returning visitors can provide valuable clues. If drop-off patterns differ on mobile devices, this may point to the layout, keyboard, validation or performance. Small segments should, however, be treated as sources of hypotheses rather than conclusive evidence that a problem affects every visitor.
This is where moderated usability testing becomes useful. When participants are asked to complete a central task, the team can observe misunderstood labels, unexpected requirements, uncertainty about the next step and information being sought in the wrong place. A recurring observation may justify a change, but it does not reveal how large the effect will be across the entire target audience.
The following observation matrix contains illustrative examples, not observations from a real client case. Its purpose is to show how qualitative findings can be connected to a measurable part of the funnel and converted into testable hypotheses.
| Task | Observed obstacle | Possible cause | Affected funnel stage | Proposed hypothesis |
|---|---|---|---|---|
| Assess whether the service is suitable | The participant looks for the scope of delivery in the footer and returns to the menu several times | Essential terms are positioned too far from the offer | Service page to form start | Explaining the scope more clearly near the main message may reduce uncertainty before the next step |
| Start a quote request | The participant hesitates when required information is requested without explanation | It is unclear why the information is needed and what happens after submission | Form start | A brief explanation of the purpose and next step may make the process easier to understand |
| Correct a validation error | The participant sees a generic error message but cannot find the incorrect field | The feedback is not linked to the relevant input field | Validation | Field-level error messages with clear correction instructions may reduce dead ends |
| Complete the request | The participant clicks several times because the confirmation is delayed or appears unclear | System status and receipt confirmation are communicated poorly | Submission and completion | Clear status information and confirmation may reduce uncertainty and duplicate submissions |
The practical sequence is often more valuable than the choice of any individual tool. Start with funnel data to locate an unusual or commercially critical drop-off. Then review relevant sessions and ask users to complete the same task to understand the likely cause. Only when the hypothesis is clear and sufficient traffic is available does A/B testing become a good way to estimate whether the change works at scale.
In some situations, the problem may be so clear that direct implementation is more reasonable than an experiment. A broken form, a misleading label or a feature that cannot be used on a common device does not need to be retained merely to create a control group. The decision should then be documented as a correction followed by monitoring, not described as a statistically proven conversion gain.

What most people get wrong about A/B testing with low traffic
The myth: All traffic can be tested
Many people believe: You can always run an A/B test as long as the website receives some visits and the test is allowed to run for long enough.
The reality: Traffic is sufficient only when the test can detect, within a reasonable time, the smallest change that would genuinely affect the business. A low baseline conversion rate, a small expected effect and split traffic may otherwise require more conversions than the website can realistically collect. A clearly defined example in Evan Miller’s Sample Size Calculator can be used to assess the baseline, minimum relevant effect, significance level and statistical power before approving the test.
The myth: Build first, calculate later
Many people believe: The most efficient approach is to build the variant immediately and determine the test duration once the experiment is underway.
The reality: Without a calculation plan, the team risks spending development time on a test that has no realistic chance of supporting a useful decision. The plan should specify the primary outcome, baseline, smallest commercially relevant effect, traffic allocation and stopping rule. When a variant requires design, development and quality assurance, a simple sample size estimate can show before development begins whether a standard A/B test is the wrong choice.
The myth: More weeks solve everything
Many people believe: If traffic is low, you can simply let the test run until the result becomes statistically significant.
The reality: A very long test is exposed to changes in campaigns, traffic mix, seasonality, products and technology. Observations made before and after a change may then describe different conditions, while repeated checking increases the risk of a misleading decision. The analysis method and stopping rule should be determined in advance, while changes to campaigns, prices, product ranges and tracking should be documented throughout the test period.
The myth: Micro-conversions can replace purchases
Many people believe: When there are too few purchases, you can optimize clicks, scrolling or button presses and assume the business results will follow.
The reality: Micro-conversions provide more data points but are only useful if their relationship with the business objective is credible. A variant may increase clicks by creating false expectations while simultaneously weakening later stages of the funnel. The entire journey should therefore be compared in a tool such as GA4 or Matomo: landing page, product or service view, process started, completed conversion and drop-off between stages.
The myth: A/B testing is the safest option
Many people believe: A/B testing always provides more reliable evidence for decisions than qualitative methods.
The reality: When conversion volume is low, funnel analysis and moderated usability testing often support a better next decision because they can reveal where users get stuck and why. They do not prove a general effect size, but they can identify specific problems that should be fixed or prioritized for later testing. Ask participants to complete the same central task while the team observes misunderstood labels, validation errors, hesitation and dead ends.
What low traffic means for your CRO work
-
Calculate test feasibility before development
Define the primary conversion goal, the smallest commercially relevant effect and the required traffic volume using a sample size calculator such as Optimizely’s Sample Size Calculator. Request results per variant, an estimated number of conversions and a timeline based on qualified traffic. If the test would need to run for an unreasonable length of time, the web agency should recommend another validation method instead of starting an experiment with no realistic ability to support a decision.
-
Prioritize clear hypotheses over minor detail tests
Spend development time on changes affecting the offer, information architecture, forms or purchase flow rather than isolated adjustments to colors and buttons. Document each hypothesis with the observed problem, proposed change, relevant target audience and business metric. The less traffic a website receives, the harder it is to justify testing effects that are likely to be small and difficult to distinguish from random variation.
-
Combine behavioral data with user insights
Use GA4 to map drop-offs in the journey and Microsoft Clarity to review session recordings and interaction patterns. Supplement this with moderated usability testing or customer interviews to understand why visitors hesitate. Qualitative observations should be treated as evidence for prioritization and hypothesis development, not as statistical proof of a particular increase in conversions.
-
Require predefined decision rules from the web agency
Define the metrics, target audience, test period, analysis method and stopping rules before the test goes live. Avoid declaring a winner based on temporary fluctuations in the experimentation tool. At the same time, verify through Google Tag Manager and GA4 DebugView that variant allocation and primary events work correctly, and investigate any unexpected imbalance between the groups before using the result to make business decisions.
Low traffic does not mean that conversion optimization is the wrong investment, but the approach must be broader than continuous A/B testing. The right next step may be an experiment, a funnel analysis, a usability test or the direct correction of a well-documented problem. An experienced team such as FLAR AB therefore starts with the decision that needs to be made and the evidence that can realistically be obtained, rather than whichever tool happens to be easiest to launch. That way, the business invests in better decisions—not merely more tests.