
Clinical trials are often designed to show that a new treatment is better than an existing treatment or control. However, superiority is not always the development question. This article looks at how equivalence trials are designed, analysed, and interpreted, including where re-randomisation may fit.
An equivalence trial is designed to demonstrate that the difference between two treatments lies within prespecified limits that represent the largest differences considered clinically acceptable. Its aim is to show, with the required statistical precision, that any difference is small enough to be considered unimportant for the trial objective.
These limits are usually described as the lower and upper equivalence margins. For a difference measure they are often written as -Δ and +Δ. If, for example, differences from -10 to +10 units have been justified as clinically unimportant, these values would define the equivalence range.
Equivalence designs can be useful when the new option may offer advantages that are not captured by the primary efficacy endpoint, such as a different route of administration, reduced treatment burden or greater convenience. The design should still be chosen because it answers the clinical question, rather than because equivalence may appear easier to demonstrate than superiority.
Superiority, non-inferiority and equivalence trial designs can use similar endpoints and comparators, but they answer different questions. Defining the objective prospectively matters because the hypothesis, margin, sample size and interpretation all depend on it.
|
Trial design |
Scientific question |
What the analysis aims to establish |
|
Superiority |
Is one treatment better than another? |
Looks for a treatment difference in a prespecified direction. |
|
Non-inferiority |
Is the new treatment no more than an acceptable amount worse? |
Rules out an unacceptable loss of effect beyond a prespecified margin. |
|
Equivalence |
Are the treatments sufficiently similar? |
Rules out unacceptable differences in both directions. |
A conventional superiority test that does not reach statistical significance has not demonstrated that the treatments are equivalent. The same non-significant result could arise because the treatments are genuinely similar, or because the study is too small or too variable to detect a meaningful difference.
It therefore cannot be established by interpreting p > 0.05 as proof of similarity.
Margins should be prespecified and justified from the clinical question, relevant evidence and knowledge of the comparator. The chosen effect scale also matters: a margin on an absolute risk-difference scale answers a different question from a margin on a ratio scale.
Clinical relevance
The margin should exclude differences that would change clinical interpretation or decision-making.
Existing evidence
Historical trials, meta-analyses or other reliable evidence may inform what treatment effect should be preserved or ruled out.
A very wide margin may make equivalence easier to demonstrate while allowing differences that are clinically important.
After the margins have been set, the analysis asks whether unacceptable differences can be ruled out at both boundaries. The Two One-Sided Tests (TOST) framework and its corresponding confidence-interval interpretation are commonly used to explain this logic.
For symmetric margins -Δ and +Δ, the null hypothesis covers treatment effects at or beyond an unacceptable boundary, whereas the equivalence alternative is that the true treatment difference lies strictly between the two limits.
Conceptually: H0: difference <= -Δ or difference >= +Δ; H1: -Δ < difference < +Δ.
TOST separates the equivalence question into two one-sided tests. One assesses whether the treatment difference is sufficiently above the lower boundary; the other assesses whether it is sufficiently below the upper boundary. Equivalence is demonstrated only when both boundary hypotheses are rejected at the prespecified one-sided significance level.
Under the conventional TOST formulation with each one-sided test at alpha = 0.05, the corresponding two-sided confidence interval is 90%. Equivalence is supported when that entire interval lies inside both prespecified equivalence limits. Other confidence levels can be appropriate where the prespecified framework or regulatory setting differs. The confidence level should therefore be interpreted within the prespecified testing framework rather than as a universal requirement for equivalence trials.
For example, with equivalence limits of -10 and +10, an estimated difference of 2 with a 90% confidence interval from -4 to +8 supports equivalence under this framework. If the same point estimate had an interval from -4 to +12, equivalence would not be demonstrated because the upper limit crosses the +10 margin.
Equivalence trials need sufficient precision to rule out clinically important differences. Sample size planning therefore needs to be tied directly to the margin, expected variability or event rates, the assumed treatment difference and the chosen power and type I error.
Wider outcome variability generally requires more participants to obtain the same precision. Narrower equivalence margins generally require more information because the confidence interval must fit inside a smaller range. Non-evaluable or missing data can reduce precision and should be anticipated in sample-size planning.
The power calculation should reflect the analysis that will be used, including the assumed true treatment difference.
An equivalence conclusion is only meaningful if the trial could have detected a clinically important difference had one existed.
ICH E10 describes assay sensitivity as the ability of a trial to distinguish effective from less effective or ineffective treatment. For non-inferiority and equivalence trials, historical evidence of sensitivity to treatment effects and appropriate conduct of the current trial are central to judging whether apparent similarity is interpretable.
Poor intervention delivery, an insensitive endpoint, or an unsuitable population can dilute genuine differences. Unlike a superiority trial, this movement towards similarity can make an incorrect equivalence conclusion more likely.
Treatment discontinuation, switching, non-adherence and major protocol deviations can all change the observed contrast between groups. Their implications depend on the clinical question and on how the treatment effect is defined.
For that reason, the usual assumption that an intention-to-treat analysis is automatically conservative should not be carried over mechanically from superiority trials. Per-protocol analyses can add useful information but are also vulnerable to selection bias. The primary and supportive analyses should be prespecified, and differences between them should be investigated rather than resolved by choosing whichever result is more favourable.
ICH E9(R1) provides a framework for defining the treatment effect of interest through an estimand and for specifying how intercurrent events, such as treatment discontinuation or rescue therapy, are handled in that clinical question. This is directly relevant to equivalence studies because different strategies can change what treatment effect is being estimated.
Missing data is distinct from intercurrent events, although the two can be related. The statistical analysis should therefore include appropriate assumptions, sensitivity analyses and data-collection plans that support the prespecified estimand.
If the allocation process creates systematic imbalance between treatment groups, the estimated treatment effect and the operating characteristics of the analysis can be affected.
Static and dynamic randomisation handle treatment allocation differently. With static methods, treatment allocations can be generated before participant characteristics are known, often with stratification. Dynamic methods, including minimisation-based approaches, allocate treatment as participants enter the study and may use selected baseline characteristics to maintain balance during recruitment.
When a complex allocation rule is used, a randomisation-based or re-randomisation analysis may be considered as a supportive way to align inference more closely with the actual allocation mechanism. The implementation must reproduce the randomisation procedure appropriately rather than simply shuffle treatment labels without regard to the design.
A superiority example helps illustrate the basic idea. Data was simulated for 200 participants allocated 1:1 to two treatments. Although the underlying treatment distributions were identical, a selected simulation seed produced a nominal p-value of 0.035 and a difference between proportions of 0.1033.
Treatment labels were then re-randomised 1,000 times and the difference between proportions recalculated. Approximately 5.2% of the simulated statistics were more extreme than the observed result. The example illustrates how a randomisation distribution can be used to assess how unusual an observed statistic is under the relevant randomisation hypothesis.

Figure 1. Re-randomised test statistic distribution. The dashed lines mark the observed statistic (±10.33 percentage points); the histogram shows the re-randomisation reference distribution.
In an equivalence trial, the important point is that the equivalence null hypothesis is not simply a sharp hypothesis of no treatment effect. It is a composite null that includes effects at or beyond the lower or upper equivalence boundary.
The worked equivalence example uses binary data from 300 participants allocated 1:1, with an illustrative true treatment difference of 10 percentage points and equivalence margins of -10 to +10 percentage points. It compares a conventional TOST analysis with a re-randomisation calculation.
Table 1. Simulation results from the equivalence example using a consistent treatment contrast.
|
Scenario |
Responders A |
Responders B |
Difference (B - A) |
90% CI |
|
True |
0.50 |
0.60 |
0.10 |
- |
|
Simulated |
0.51 |
0.54 |
0.03 |
-0.037 to 0.097 |
Note: Results are expressed using consistent B - A contrast. The TOST p-value is 0.0429; with one-sided alpha = 0.05 this corresponds to the 90% confidence-interval framework described above.
A randomisation procedure developed for a no-effect superiority hypothesis should not automatically be interpreted as a valid test of an equivalence boundary hypothesis. Any randomisation-based equivalence analysis needs a null hypothesis, test statistic and re-randomisation mechanisms that are coherent with one another.
Randomisation-based methods can be useful in appropriately justified settings, including as prespecified supportive analyses. They should complement, rather than replace, the primary equivalence analysis.
Regulatory expectations depend on the product, indication, development programme and jurisdiction. The principles below provide context for planning and interpretation, but they should not be treated as a single universal rulebook for every equivalence study.
In November 2025, the European Medicines Agency published the draft Guideline on non-inferiority and equivalence comparisons in clinical trials (EMA/301654/2025). The consultation ran from 13 November 2025 to 31 May 2026 and is now closed; as of September 2026, the document is still listed by EMA as a draft.
The guidance covers planning, conduct, analysis and interpretation, including trial objectives, estimands, assay sensitivity, trial quality, margin selection and statistical considerations. Because it remains draft guidance, its provisions should be described as proposed rather than as adopted requirements.
ICH E10 remains a key source for active-control trials, while ICH E9(R1) provides the estimand framework.
FDA's final 2016 guidance on Non-Inferiority Clinical Trials is specific to non-inferiority rather than a general equivalence guideline. It is nevertheless relevant background for active-control studies because it addresses interpretability, margin selection and testing of the non-inferiority hypothesis. Equivalence-specific conclusions should not be inferred from NI guidance without considering the relevant development context.
In conclusion, re-randomisation remains a useful specialist consideration, but its interpretation must match the equivalence hypothesis rather than simply reuse the logic of a superiority permutation test.
If prespecified equivalence limits are -10 to +10 units and the relevant confidence interval for the treatment difference is -4 to +8, equivalence can be supported because the whole interval lies within both limits. If the interval extends beyond either boundary, equivalence is not demonstrated.
Equivalence testing assesses whether a treatment, intervention or method differs from its comparator by less than prespecified acceptable limits. It therefore asks whether meaningful differences can be ruled out, rather than whether the treatments are the same.
A superiority trial seeks evidence that one treatment performs better than another. A non-inferiority trial instead seeks to rule out a prespecified unacceptable loss of effect for the new treatment.
Quanticate's statistical consultants are among the leaders in their respective areas enabling the client to have the ability to choose expertise from a range of consultants to match their needs. If you have a need for these types of services please request a consultation and a member of our Business Development team will be in touch with you shortly.
Bring your drugs to market with fast and reliable access to experts from one of the world’s largest global biometric Clinical Research Organizations.
© 2026 Quanticate