Fragility metrics provide a general framework for quantifying how robust or fragile statistical results are to small changes in the underlying data. The p–fr–nb framework described here is designed to apply across a wide range of study designs, including binary, continuous, and time-to-event outcomes.
This documentation covers classical fragility measures such as the Fragility Index (FI), Fragility Quotient (FQ), and Marginal Fragility Quotient (MFQ), as well as more general extensions including the GFI and GFQ. It also outlines fragility methods for ANOVA, regression models, and other settings beyond simple two-arm 2×2 tables.
The content below is synced automatically from the official
FRAGILITY_METRICS.md file in the
fragility-metrics GitHub repository
.
Thomas F. Heston Department of Family Medicine, University of Washington, Seattle, WA, USA Department of Medical Education and Clinical Sciences, Washington State University, Spokane, WA, USA ORCID: 0000-0002-5655-2512 Version: 13.3.0
A p-value of 0.049 and a p-value of 0.0001 are both reported as 'statistically significant'—but they represent vastly different levels of evidence quality. The p–fr–nb framework fixes this. Instead of reporting p-values alone ("partial evidence"), we propose complete statistical evidence, defined as the triplet p–fr–nb: the p-value (significance), a native fragility quotient fr (classification stability), and a neutrality-boundary robustness metric nb (distance from therapeutic neutrality). Fragility (fr) is first quantified by native fragility quotients that measure the proportion of relevant data (or SE-scale shift) required to flip significance classification within a given design, with primary metrics MFQ, GFQ (gold standard for r×c and multinomial), DFQ (diagnostic benchmarks), BFQ (single-arm benchmarks), CFQ (continuous outcomes via Welch t-geometry), PFI (fixed-margin designs), ANOVA-FQ (multi-group continuous outcomes), ZFQ (the Fisher-z Fragility Quotient; correlations), OFQ (ordinal outcomes via Wilcoxon-Mann-Whitney z-statistic), and SFQ (survival outcomes via Cox regression z-statistic). fr is the native fragility quotient for the design at hand (MFQ, GFQ, CFQ, …), computed directly from the observed data; high fr indicates a stable classification, low fr indicates fragility. Native quotients are not numerically comparable across designs; a cross-design percentile scale is a deferred extension (see Part I). Robustness (nb) quantifies geometric distance from therapeutic neutrality via the Neutrality Boundary Framework (NBF), with primary metrics RQ (independent-sample binary/multinomial), MHQ (matched-pair/fixed-margin designs), DNB (diagnostic odds ratio), Proportion-NBF (single-arm benchmarks and agreement vs chance), MeCI (continuous means), DTI (correlation), ANOVAη² (multi-group), ORQ (ordinal outcomes), and SRQ (survival outcomes). All metrics use only observed counts or published summary statistics; no raw data, simulation, reconstruction, or covariate models permitted. Fragility always measures classification stability (high fr is desirable when the p-value supports the claim). Robustness interpretation is claim-dependent: high nb supports "effect exists" claims, undermines "no effect" claims. This document finalizes the integration of continuous-outcome measures (CFQ, MeCI, ANOVA-FQ, ZFQ), single-arm benchmark measures (BFQ + Proportion-NBF), and the unified fr/nb notation, providing a complete evidence-quality system applicable to nearly every standard study design with minimal assumptions. We define "complete statistical evidence" as the p–fr–nb triplet: p for significance, fr for fragility, and nb for robustness, replacing partial evidence based only on statistical significance or nonsignificance.
Whereas p < 0.05 establishes statistical significance, a concordant p–fr–nb triplet (low p + high fr + high nb) establishes convincing, complete statistical evidence. Traditional practice reports only the p-value for a given analysis, which we term partial evidence: it addresses compatibility with the null but not the stability of that decision or the distance from therapeutic neutrality. We define complete statistical evidence as the triplet (p, fr, nb), where p quantifies statistical significance; fr is the native fragility quotient for the design (MFQ, GFQ, CFQ, etc.), a 0–1 measure of the proportion of data or SE-shift needed to flip significance — high fr means the classification is stable, low fr means fragile; and nb is a 0–1 robustness metric measuring geometric distance from the neutrality boundary. Only when all three dimensions align (low p, high fr, high nb for "effect exists" claims) do we regard the statistical evidence as complete in the sense of being decision-ready and replication-ready. Reporting only p-values yields partial evidence, because it ignores both the stability of the conclusion (fragility) and the distance from neutrality (robustness). The p–fr–nb triplet restores these missing dimensions and constitutes complete statistical evidence for a result.
You have a 2×2 table from a binary outcome trial:
| Events | Non-events | |
|---|---|---|
| Treatment | 13 | 87 |
| Control | 25 | 75 |
| Calculate: |
Modern evidence assessment rests on three complementary statistical dimensions plus clinical effect size:
| Secret question from every skeptical reader | p alone answers | p–fr–nb triplet answers |
|---|---|---|
| Significant? | Yes | Yes (p) |
| Flippable by a few outcome changes or dropouts? | ? | Yes — fr quantifies the fragility of the p-value |
| Real separation from zero, or just lucky sampling noise that barely hit p<0.05? | ? | Yes — nb quantifies distance from the neutrality boundary |
| Only when all three dimensions align strongly (low p + high fr + high nb) do we have truly compelling, replication-ready evidence that an intervention works. |
Probability (p-value): the p-value quantifies the compatibility of the observed data with the null hypothesis (no effect). Lower p-values indicate stronger evidence against the null hypothesis. Conventional threshold: p < 0.05 for "statistically significant." Fragility (fr): the fragility summary statistic, fr, measures the stability of the significance classification. A high fr indicates stability, i.e., it takes a significant shift in outcomes to flip significance. A low fr indicates fragility, i.e., it takes only a slight change in outcomes to flip significance. Fragility quantifies the minimal perturbation to the data required to reverse the p-value decision. fr ∈ [0,1] is the native fragility quotient for the design (e.g., MFQ, GFQ, CFQ), computed directly from the observed data. Robustness (nb): The robustness summary statistic, nb, measures how far the observed result sits from therapeutic neutrality (no effect), expressed as a bounded, sign-agnostic standardized effect magnitude on a 0–1 scale. nb ∈ [0,1] where high nb = far from neutrality and low nb = near neutrality. nb is a property of the point estimate: it is computed from the observed effect magnitude, not from its precision — sampling uncertainty is carried by p and by fragility, not by nb. This separation is deliberate: it is what lets the triplet distinguish a large-but-imprecise effect (high nb, fragile) from a genuinely null one (low nb), and it is what the founding metric RQ already does (RQ is scale-invariant — multiplying every cell of a 2×2 by a constant leaves it unchanged). Here "robustness" denotes distance from therapeutic neutrality — a standardized effect magnitude — and not the classical statistical sense of insensitivity to modeling assumptions or outliers. A high nb indicates the result is far from neutrality; a low nb indicates it is statistically close to neutrality (which is not, by itself, affirmative evidence that no effect exists). nb is comparable across trials within a design; native nb values are not equivalent across designs (empirically they diverge), so cross-design nb supports a common interpretive language (weak/moderate/strong), not numerical equivalence. Effect size: the magnitude of the observed effect. Two magnitudes matter, and the framework now separates them cleanly. The relative, standardized magnitude — how far the effect sits from no effect on a common 0–1 scale — is captured by nb itself: SRQ reparametrizes ln(HR), DTI reparametrizes atanh(r), and the rest of the transform family behaves the same way. The absolute magnitude — mean difference in native units, absolute risk reduction, number needed to treat, months of survival gained — is not recoverable from nb, and it is the quantity the fourth element supplies. Like nb, the absolute effect size is a property of the point estimate, independent of statistical significance and sampling uncertainty. The split is therefore relative magnitude (nb, inside the triplet) versus absolute magnitude (effect size, the fourth element), not "should I believe it" versus "how much better," since the triplet already speaks to relative magnitude through nb. A large relative effect can still be a trivial absolute one: halving risk from 2% to 1% yields a healthy nb but a number needed to treat of 100. Complete evidence therefore pairs the triplet with the absolute effect size; clinical decisions require both. Partial Evidence: reporting of p-values alone or p-values with 95% CIs only constitutes "partial evidence." Complete Statistical Evidence: a result is considered to have complete statistical evidence only when all three dimensions of the p–fr–nb triplet are reported together: significance (p-value), fragility (fr), and robustness (nb). Traditional reporting of the duplet p-values with 95% confidence intervals (CI) constitutes "partial evidence." The p-value addresses only compatibility with the null hypothesis, while the 95% CI quantifies precision and effect size, but does not directly measure classification stability or normalized strength of evidence for a non-zero effect. The CI tells you the range of plausible effect sizes but not how many outcome changes would flip statistical significance (fragility); nor does it provide a standardized measure of how strong the evidence is that a real, non-zero effect exists (this is what robustness quantifies on a 0–1 scale). Complete evidence requires assessing all three dimensions to determine whether a finding is decision-ready and replication-ready. Recommended reporting thus includes complete statistical evidence (p–fr–nb) plus the non-statistical (but critical) quantity, effect size. Complete Evidence: The p–fr–nb triplet plus effect size (the quartet). Complete evidence requires both the inferential assessment (is the finding real, stable, and separated from null?) and the magnitude assessment (how large is the effect?).
Native fragility quotients (fr ∈ [0,1])
Cross-design normalization (deferred). Native quotients are not numerically comparable across designs (binary medians run lower than continuous). A percentile-normalized universal scale is a planned extension pending published reference distributions; until then, fr is interpreted within its design family.
Claiming an effect exists (p ≤ 0.05): Prefer low p, high fr, high nb. Only when all three align strongly do we have truly compelling evidence.
Claiming no effect (p > 0.05): Prefer high p, high fr, low nb.
The measurements are universal; the interpretation is claim-dependent.
In most common trial designs, the framework provides paired metrics (both fr and nb) so stability and distance from neutrality can be evaluated together (e.g., MFQ + RQ, PFI + RQ, DFQ + DNB, CFQ + MeCI, ANOVA-FQ + ANOVAη²).
Native fragility quotients (MFQ, GFQ, CFQ, SFQ, ZFQ, OFQ, ANOVA-FQ, PFI, DFQ, BFQ) provide direct physical interpretation within each study design, but their raw values have not been validated to be comparable across designs. A universal cross-design scale via percentile normalization against reference distributions is a possibility. Pending further work, fr is the native quotient, interpreted within its design family.
Measurement ≠ Interpretation Robustness (nb) has opposite implications depending on the claim being made:
| Claim | High Robustness (far from neutral) | Low Robustness (near neutral) |
|---|---|---|
| "Effect exists" (p ≤ 0.05) | Supports the claim | Undermines the claim |
| "No effect" (p > 0.05) | Undermines the claim | Supports the claim |
| Fragility (fr) behaves differently: |
| Metric | Type | Scale | Primary/Secondary | Formula (core) | Purpose |
|---|---|---|---|---|---|
| FQ | Fragility | 0–1 | LEGACY | FI / N | Proportion to flip (classic, total N) |
| MFQ | Fragility | 0–1 | PRIMARY | FI / n_mod | Proportion to flip (arm-specific) |
| GFQ | Fragility | 0–1 | PRIMARY | GFI / N | Proportion to flip (global, r×c) |
| DFQ | Fragility | 0–1 | PRIMARY | DFI / n_relevant | Proportion to flip (diagnostic) |
| BFQ | Fragility | 0–1 | PRIMARY | BFI / n_relevant (n_relevant = n) | Proportion to flip (single-arm vs benchmark) |
| CFQ | Fragility | 0–1 | PRIMARY | ||T| − t*| / (1 + ||T| − t*|) | SE-scaled distance to p = 0.05 (continuous) |
| PFI | Fragility | 0–1 | PRIMARY | 4 × |x| / N (x = fixed-margin path shift) | Independent-sample 2x2 sub-integer fragility |
| RQ | Robustness | 0–1 | PRIMARY | Σ|O − E| / [2N(m − 1)/m], m = min(r, c); for any 2×2 this equals Σ|O − E| / N = |ad − bc| / (N²/4) | Distance from independence |
| sRQ | Robustness | −1–+1 | PRIMARY | 4(ad − bc) / N² (signed RQ; |sRQ| = RQ) | Signed distance from independence |
| wsRQ | Robustness | −1–+1 | PRIMARY (meta) | Σ(sRQ_i · w_i), w_i = N_i/ΣN | Pooled signed meta-analytic robustness |
| wGFQ | Fragility | 0–1 | PRIMARY (meta) | Σ(GFQ_i · w_i), w_i = N_i/ΣN | Pooled meta-analytic fragility |
| MHQ | Robustness | 0–1 | PRIMARY (matched) | |b − c| / (b + c) or 0 if b + c = 0 | Distance from marginal homogeneity |
| DNB | Robustness | 0–1 | PRIMARY | |ln(DOR)| / (1+|ln(DOR)|) | Diagnostic distance from neutrality |
| Proportion-NBF | Robustness | 0–1 | PRIMARY | |p̂ − p₀| / (|p̂ − p₀| + √[p₀(1 − p₀)/n_relevant]) | Single-arm distance from benchmark / chance agreement |
| MeCI | Robustness | 0–1 | PRIMARY | d / (1 + d) where d = min(|μ₁−c|, |μ₂−c|) / √(s₁²+s₂²), c = (s₁μ₂+s₂μ₁)/(s₁+s₂) | Continuous distance from neutrality |
| DTI | Robustness | 0–1 | PRIMARY | |atanh(r)| / (1 + |atanh(r)|) | Correlation distance from independence |
| ZFQ | Fragility | 0–1 | PRIMARY | |Z − 1.96| / (1 + |Z − 1.96|) where Z = |atanh(r)|√(n−3), n > 3 | Correlation classification stability (Fisher-z) |
| OFQ | Fragility | 0–1 | PRIMARY | ||z_WMW| − 1.96| / (1 + ||z_WMW| − 1.96|) | SE-scaled distance to p = 0.05 (ordinal) |
| ORQ | Robustness | 0–1 | PRIMARY | |ln(gOR)| / (1 + |ln(gOR)|) | Distance from neutrality (ordinal) |
| ANOVA-FQ | Fragility | 0–1 | PRIMARY (k≥2) | |√F − √F| / (1 + |√F − √F|) | Stability of F-classification |
| ANOVAη² | Robustness | 0–1 | PRIMARY | df_b·F / (df_b·F + df_w) | Distance from equality of means |
| FI | Count | 0–N | Secondary | Toggle count (classic) | Raw fragility count (binary) |
| SFI | Count | 0–N | Secondary | Toggle count (standardized) | Label-invariant count |
| GFI | Count | 0–N | Secondary | Move count (global) | Path-independent count |
| DFI | Count | 0–N | Secondary | Toggle count vs benchmark | Diagnostic count |
| NDI | Count (robustness) | 0–N/4 | Secondary | round(N·RQ/4) = round(|ad − bc|/N), clamped to reachability | Coupled fixed-margin moves to neutrality (RR = 1) |
| CFS | Distance | 0–∞ | Secondary | ||T| − t*| | SE-unit distance to p = 0.05 (continuous) |
| SFM | Scaling | > 1 | Secondary | Factor k > 1 to flip (×k flips nonsignificant → significant; ÷k flips significant → nonsignificant) | Sample size fragility multiplier |
| UFI | Unit | >0 | LEGACY | N/(n₁n₂) or 1/max(n₁, n₂) or 1/N | Step-size definitions (fixed-margin unit size) |
| SFQ | Fragility | 0–1 | PRIMARY | ||z_HR| − 1.96| / (1 + ||z_HR| − 1.96|) | SE-scaled distance to p = 0.05 (survival) |
| SRQ | Robustness | 0–1 | PRIMARY | |ln(HR)| / (1 + |ln(HR)|) | Distance from neutrality (survival) |
| t* is the critical value from the t-distribution. | |||||
| F* is the critical F value at α = 0.05 for the reported df. | |||||
| m = min(r, c). The denominator 2N(m − 1)/m is the maximum of Σ | O − E | , attained under perfect association; for 2×2 it equals N, so 2×2 values are unchanged. |
Fragility quotients measure the proportion of the sample (binary/diagnostic) or the proportion of SE-scale movement (continuous) required to flip statistical significance. All primary native fragility quotients range 0–1 (MFQ, GFQ, CFQ, SFQ, etc.). fr is the native fragility quotient for the design at hand, computed directly from the observed data. Cross-design percentile normalization is a deferred extension (see Part I), not part of the operational framework.
Application: Legacy metric for any binary outcome 2×2 table (independent samples) using total N denominator. Definition: Proportion of the total sample that must toggle to flip significance Formula: FQ = FI/N Range: 0 to 1 Interpretation: fr = FQ (the native quotient). For example, FQ = 0.02 means 2% of sample outcomes must change to flip statistical significance. Advantages: use for historical comparison with studies that used FQ Base metric: FI (classic fragility index) NBF pair: RQ Note: FQ is a legacy metric and is not recommended as the primary fr metric for 2-arm binary outcome studies. GFQ is preferred (path-independent and label-invariant); when large N makes GFI computation intractable, MFQ is the fallback, which is allocation-fair (denominating against the arm actually modified) and label-resistant.
Application: Fallback for independent-sample 2×2 binary outcome trials (any allocation ratio) when large sample size (≈5000+) makes GFI computation intractable, or when compatibility with the classic FI count is required. Definition: Proportion of the arm that was actually modified in the classic fragility index procedure required to flip statistical significance. Formula: MFQ = FI / n_mod, where n_mod = sample size of the arm subjected to toggling in the standard FI calculation (i.e., the arm with fewer events; if tied, the smaller arm). Range: 0 to 1 Interpretation: fr = MFQ. Example: MFQ = 0.05 means 5% of patients in the arm that was toggled would need to switch outcome to flip significance. Advantages: allocation-resistant: minimizes the distortion caused by unequal study allocation; the MFQ is label-resistant but not label independent like the GFQ is. Base metric: Heston FI NBF pair: RQ Note: For 2×2 binary outcome tables, GFQ is the recommended default (§3.3); use MFQ when GFI is computationally intractable at large N, or when comparability with the universally recognized FI count is needed. MFQ remains allocation-fair by denominating against the arm that actually needs to change. When MFQ is used, pair it as MFQ + RQ.
Application: Recommended default and gold standard for any r×c contingency table or multinomial outcomes (independent samples), including 2×2 binary outcome trials. Definition: Proportion of sample involved in minimal cell moves Formula: GFQ = GFI/N Range: 0 to 1 Interpretation: For GFQ, fr = GFQ (e.g., GFQ = 0.03 means 3% of the sample must be reallocated to flip statistical significance). Advantages: Path-independent, applies to any r×c table Base metric: GFI (global fragility index) NBF pair: RQ Note: Considered the gold standard for binary and multinomial fragility assessment because of its path-independence and label-invariance, and it is the recommended default for 2×2 binary outcome tables. For large sample sizes (≈5000+), computing GFI becomes computationally intractable, and MFQ becomes the practical fallback for two-arm binary outcome studies (§3.2). The GFQ complements RQ, which measures robustness (distance from independence). Both should be reported together — GFQ + RQ is the standard pair for binary and multinomial outcomes. Both the GFI and the GFQ show label-invariance.
Application: Diagnostic metrics from full 2×2 table (TP, FN, FP, TN) with ground truth. Definition: Proportion of the relevant subset of observations that must toggle to change the diagnostic classification (below vs not below benchmark) using a one-sided exact binomial test against p₀, paired with DNB. Formula: DFQ = DFI / n_relevant, where n_relevant depends on the metric:
Rationale: We cannot compare a test's PPV/NPV/ACCURACY across studies unless the underlying prevalence is normalized. Since clinical testing is most helpful when diagnostic uncertainty is high (pretest ≈ 50%), prevalence is set to 50%. Procedure:
Application: Single-arm benchmark tests and agreement vs benchmark when only (k, n, p₀) are available. Typical use cases: a) Single-arm response rate vs a benchmark proportion p₀ b) Single-rater agreement vs chance (p₀ = 0.5) or another target benchmark when the full TP/FN/FP/TN table is not available Definition: Proportion of observations in a single-arm proportion that must toggle (success ↔ failure) to change the benchmark classification under a one-sided exact binomial test. Formula: BFQ = BFI / n_relevant, where:
Application: independent-sample binary 2×2 trials (continuous analog of FI/SFI/UFI for sub-integer fragility resolution). Definition: Proportion of one balanced cell's expected count that must be reallocated along a fixed-margin perturbation path to flip the Pearson chi-square significance classification in an independent-sample 2×2 design. Formula: PFI = |x| / (N/4) = 4|x| / N, where x is the smallest margin-preserving change in the four cells, applied along the path (a+x, b−x, c−x, d+x), that flips the two-sided Pearson chi-square decision across the α = 0.05 boundary (sign ignored), and N/4 is the expected count per cell in a perfectly balanced 2×2 under independence. For reporting, PFI is conventionally multiplied by 100 and expressed as a percent (e.g., PFI = 0.02 is reported as 2%). Range: 0 to 1 (native quotient; reported as 0% to 100%). Bounded by construction for significance flips: along the fixed-margin path the cross-product difference is linear, Δ(x) = (ad − bc) + xN, so the independence point (Δ = 0) sits at |x| = |ad − bc| / N ≤ N/4 (since |ad − bc| ≤ N²/4). The two-sided chi-square boundary is crossed at or before independence, hence x_flip ≤ N/4 and PFI = 4|x_flip|/N ≤ 1. Boundary-limited cases (below) are capped at 1.0. Significance test: Pearson chi-square, without Yates correction. Fisher's exact does not admit fractional counts, so it cannot resolve sub-integer flips; Yates correction is omitted because it is designed for integer counts, adds conservative bias, and disrupts the smoothness of χ² that continuous perturbation requires. Interpretation: fr = PFI. Lower PFI indicates greater fragility. PFI = 0.02 (2%) means a 2% proportional shift in outcome distribution, relative to the balanced-cell expected count, would reverse significance. Provisional guidance (subject to empirical testing): PFI < 0.05 (5%) is concerning for fragility; PFI > 0.10 (10%) suggests relatively stable findings. Advantages: a) detects sub-integer fragility that integer metrics (FI, SFI) cannot represent; b) label-agnostic — no determination of minority arm or event coding required; c) margin-preserving, so perturbations reflect proportional redistribution within the observed trial structure; d) continuous scale aligns naturally with the framework's other continuous fragility metrics (CFQ, SFQ, ZFQ, OFQ) and supports smooth percentile normalization to fr. Base metric: x (minimal margin-preserving cell perturbation along the fixed-margin path) NBF pair: RQ Note: The fixed-margin path is a sensitivity construct, not a claim that the trial had fixed margins by design. It holds row sums (arm sizes) and column sums (total events) constant so that any admissible change is a coupled diagonal exchange, isolating sensitivity to proportional redistribution. The continuous (chi-square) criterion is required for sub-integer resolution; PFI is a descriptive evidence-quality metric, not an inferential test, so the perturbation geometry and the significance criterion need not share a single likelihood. For strictly matched-pair or crossover designs in which both margins are fixed by the design itself, McNemar-based fragility (with MHQ robustness) is the design-coherent choice; PFI as defined here targets the independent-sample case.
Sections §3.7–3.11 are one construct — the bounded distance from the test statistic to its α = 0.05 critical value, fr = δ/(1 + δ) with δ = |stat − crit| — instantiated per design: continuous (CFQ), multi-group (ANOVA-FQ), correlation (ZFQ), ordinal (OFQ), and survival (SFQ). The umbrella is conceptual; the per-design instances below are what you compute.
Application: trials comparing two continuous outcomes where m₁, m₂, s₁, s₂, n₁, n₂ are all known. Definition: Proportion of an SE-scaled shift in the estimated mean difference required to flip statistical significance in a two-sample continuous comparison (Welch t-test). Formula: Let m₁, m₂ = observed group means s₁, s₂ = observed standard deviations n₁, n₂ = group sample sizes. v₁ = s₁²/n₁ (variance of mean 1), v₂ = s₂²/n₂ (variance of mean 2), SE_diff = √(v₁ + v₂) (this is the Continuous Fragility Unit, CFU), θ̂ = m₁ − m₂ (observed mean difference), T = θ̂ / SE_diff (observed Welch t-statistic), df = (v₁ + v₂)² / (v₁²/(n₁−1) + v₂²/(n₂−1)), (Welch–Satterthwaite degrees of freedom), t = t₀.₉₇₅,df (two-sided critical value at α = 0.05). Then: Continuous Fragility Score (distance in SE units to the p = 0.05 boundary): CFS = | |T| − t |. Continuous Fragility Quotient: CFQ = CFS / (1 + CFS). Range: 0 to 1. Interpretation: fr = CFQ. For example, CFQ = 0.12 means the observed t-statistic lies relatively close to the p = 0.05 boundary on the CFQ scale; smaller values indicate a more fragile significance classification, larger values a more stable one. Advantages: Works directly from reported summary statistics (m₁, m₂, s₁, s₂, n₁, n₂). No raw data required. No simulated data or distributional reconstruction. Correctly respects Welch's variance structure and degrees of freedom. Provides a continuous-outcome analogue of MFQ/GFQ. Base metric: CFS = continuous fragility score (SE-unit distance between |T| and t*). NBF pair: MeCI Note: CFQ assesses fragility (stability of significance). It complements MeCI, which assesses robustness (distance from neutrality). Both should be reported for continuous outcomes. CI-only implementation note: When only a 95% CI for the mean difference is available, T and SE are reconstructed using a large-sample z-based approximation (t ≈ 1.96). Under this approximation, p and CFS/CFQ values are asymptotic. MeCI cannot be computed from a mean-difference CI alone because the crossover-point formula requires separate group means and SDs; in CI-only scenarios, report CFQ alone or reconstruct group-level summary statistics where possible. Note on Paired/Matched Data: For crossover trials, pre-post measurements, or matched pairs, the CFQ formula applies directly using the paired t-statistic:
Application: One-way ANOVA with k ≥ 2 independent groups (continuous outcome, equal or unequal variances assumed by the reported F-test). Definition: Bounded (0–1) distance in √F space from the α = 0.05 classification boundary. At k = 2, √F space coincides exactly with the SE-scaled t space of CFQ; for k > 2 it serves as the analogous distance scale. Formula: ANOVA-FS = |√F − √F| ANOVA-FQ = ANOVA-FS / (1 + ANOVA-FS) where F is the critical F value at α = 0.05 for the reported (df_b, df_w). Range: 0 to 1 Interpretation: fr = ANOVA-FQ (higher = more stable classification) Advantages:
Application: Pearson or Spearman correlation reported as (r, n), with n > 3 and |r| < 1. Definition: Symmetric fragility metric measuring distance from the α = 0.05 classification boundary in Fisher-z space. Formula: Let z_r = atanh(r), Z = |z_r|√(n−3), Z_crit = 1.96 D = |Z − Z_crit| ZFQ = D / (1 + D) For Spearman correlations, the Fisher-z variance is inflated; use the Fieller/Bonett–Wright adjustment Z = |atanh(r_s)| · √((n−3)/1.06) in place of the Pearson form. Domain: Requires n > 3 (√(n−3) undefined otherwise) and |r| < 1 (atanh diverges at ±1). Range: 0 to 1 Interpretation: fr = ZFQ. Higher values indicate more stable classification (significant or non-significant). Advantages: Sample-size dependent, symmetric around decision boundary, structurally identical to CFQ/ANOVA-FQ Base metric: D (distance in Fisher-z test statistic units) NBF pair: DTI Note: Completes the p–fr–nb triplet for correlation analyses. ZFQ measures classification stability; DTI measures robustness (distance from independence). Both required for complete correlation evidence assessment.
Application: Ordinal outcomes analyzed via Wilcoxon-Mann-Whitney test, proportional odds models, or ordinal logistic regression (e.g., modified Rankin Scale, NIHSS, pain scales, functional status scores). Definition: Proportion of SE-scaled shift in the ordinal test z-statistic required to flip statistical significance for ordinal outcomes. Formula: Let gOR = generalized odds ratio (common odds ratio from proportional odds model), with 95% CI [CI_lower, CI_upper], where gOR, CI_lower, and CI_upper are all > 0. Calculate:
Application: Time-to-event outcomes analyzed via Cox regression (e.g., overall survival, progression-free survival, time to heart failure hospitalization). Definition: Proportion of SE-scaled shift in the Cox regression z-statistic required to flip statistical significance for survival outcomes. Formula: Let HR = hazard ratio from Cox regression, with 95% CI [CI_lower, CI_upper]. Calculate:
Range: 0 to 1 Interpretation: fr = SFQ. Example: SFQ = 0.15 means the z-statistic is relatively close to the p = 0.05 boundary on the SFQ scale; higher values indicate more stable significance classification. Advantages: Works directly from reported HR and 95% CI (standard reporting for survival outcomes). No raw survival data, Kaplan-Meier curves, or censoring information required. Extends the fragility framework to time-to-event analysis. Base metric: | |z_HR| − 1.96 | (raw distance in z-statistic units to significance boundary) NBF pair: SRQ Note: SFQ assesses fragility (stability of significance classification) for survival outcomes. It complements SRQ, which measures robustness (distance from neutrality). Both should be reported together for time-to-event studies. Structurally identical to CFQ (continuous), ANOVA-FQ (multi-group), ZFQ (correlation), and OFQ (ordinal).
NBF metrics quantify geometric distance from neutrality—where treatment equals control (e.g. RR=1, Δ=0, r=0, DOR=1).
General NBF Formula: NBF = |T − T₀| / (|T − T₀| + S), where T = statistic, T₀ = neutral value, and S = a fixed, sample-independent scale parameter (commonly S = 1, giving the x/(1+x) bounding map). S is not a standard error: nb depends on the effect magnitude, not its precision.
Universal property: all NBF metrics → [0,1].
Interpretation: 0 = at neutrality, 1 = maximally separated.
Distance-from-neutrality transform family. DTI, ORQ, SRQ, and DNB share one construct — a bounded, sign-agnostic distance from neutrality of the form |g(θ)|/(1 + |g(θ)|), where g is a variance-stabilizing transform of the effect estimate θ (atanh for correlation r; ln for the ratio metrics gOR, HR, DOR). One idea, instantiated per design; each keeps its own θ, and each is computed from the effect estimate alone, not its precision.
Application: Independent-sample binary or multinomial outcome tables (2×2 or r×c) assessing separation from independence (treatment vs control or multi-arm studies).
Definition: NBF-based robustness metric measuring geometric distance from independence in binary or multinomial outcome tables.
Formula (general): RQ = Σ|O − E| / [2N(m − 1)/m] for any r×c table, where O are observed counts, E are expected counts under independence, m = min(r, c), and the denominator 2N(m − 1)/m is the maximum attainable value of Σ|O − E| (attained under perfect association). This normalization guarantees RQ ∈ [0, 1] for every r×c table.
Special-case shortcut (for any 2×2): since m = 2 makes the denominator equal N, RQ = Σ|O − E| / N = |ad − bc| / (N²/4) for any 2×2 table.
Range: 0 to 1.
Interpretation: nb = RQ. For example, nb = 0.20 means the data are moderately separated from independence.
Neutrality: Independence of variables (e.g. ad = bc for 2×2)
Pairs with: FQ, MFQ, GFQ, PFI
Count counterpart: NDI (Part V) — the integer count of coupled fixed-margin moves to neutrality. For any 2×2 table, NDI = round(N·RQ/4) (clamped to reachability). NDI and RQ quantify the same underlying cross-product distance from neutrality in different units: NDI expresses that distance as an integer number of coupled fixed-margin moves, whereas RQ expresses it as a normalized 0–1 geometric distance. NDI:RQ therefore forms an Index:Quotient pairing by structural analogy to GFI:GFQ, but not by direct division: RQ is not NDI/N.
Note: Standard robustness measure for independent-sample binary and multinomial outcomes.
Application: Independent-sample 2×2 outcome tables where the direction of separation from independence (which arm is favored), not only its magnitude, must be retained — required for signed meta-analytic pooling (wsRQ). Definition: The canonical RQ with the absolute value removed, so the metric carries both distance from neutrality and direction. Formula: sRQ = 4(ad − bc) / N² Range: −1 to +1 (|sRQ| = RQ). Direction convention: For a 2×2 table with rows = arms (a, b = events, non-events of row 1; c, d = events, non-events of row 2) and the first row taken as the experimental/treatment arm, the sign of (ad − bc) reflects table orientation and must be declared per analysis. In the Zuin et al. PE-thrombolysis reanalysis convention used in canon, negative sRQ favors the experimental/treatment arm (fibrinolysis); positive sRQ favors the control/harm direction. State the orientation explicitly whenever sRQ is reported. Neutrality: Independence (ad = bc → sRQ = 0). Pairs with: MFQ, GFQ (fragility); aggregates to wsRQ (meta-analytic). Note: sRQ is a robustness metric — it measures signed distance from therapeutic neutrality, never classification stability. |sRQ| recovers the unsigned canonical RQ of §4.1. Direction is a property of the table orientation, not of the metric; the convention must be fixed before pooling.
Application: Diagnostic accuracy studies with full 2×2 tables (TP, FN, FP, TN) and ground truth, including sensitivity, specificity, PPV, NPV, and accuracy (after prevalence normalization when needed).
Definition: NBF-based robustness metric measuring the diagnostic odds ratio's (DOR) distance from neutrality.
Formula: DNB = |ln(DOR)| / (1 + |ln(DOR)|)
where:
Range: 0 to 1.
Interpretation: For DNB, nb = DNB. For example, nb = 0.35 means the diagnostic odds ratio is clearly separated from no-discrimination.
Neutrality: DOR = 1 (test no better than chance).
Pairs with: DFQ
Note: DNB uses the uncertainty-free form |ln(DOR)|/(1 + |ln(DOR)|), placing it in the distance-from-neutrality transform family alongside DTI, ORQ, and SRQ (Part XI). The prior form |ln(DOR)|/(|ln(DOR)| + SE) was a monotone transform of the Wald z-statistic and duplicated the significance axis; the current form is the pure standardized effect magnitude, with precision carried by p and fr. When any cell is zero, ln(DOR) is undefined; apply a 0.5 continuity correction to all four cells before computing DNB. Primary robustness measure for diagnostic tests using a full 2×2 table. Apply after prevalence normalization for PPV/NPV and accuracy where appropriate (see §3.4.1). For single-arm benchmark analyses based only on (k, n, p₀), use Proportion-NBF instead of DNB.
Application: Two-group independent continuous outcome studies (e.g., trials reporting group means and standard deviations for each arm; sample sizes are not required for MeCI but are needed for the paired CFQ calculation).
Definition: NBF-based robustness metric measuring distributional distinguishability for continuous outcomes, based on the equal-standardized-distance point of the two sample distributions — the point equidistant from the two group means in standard-deviation units, which coincides with the density crossover point of the two distributions only when s₁ = s₂. MeCI is p-value independent and sample-size independent.
Formula: Let μ₁, μ₂ be the observed group means and s₁, s₂ their standard deviations.
Interpretation: For MeCI, nb = MeCI. Low MeCI values indicate near-equivalence between groups (high population overlap, weak separation); high MeCI values indicate robust separation (minimal overlap, strong distinguishability).
Neutrality: μ₁ = μ₂ (both means coincide with c; d = 0, MeCI = 0).
Pairs with: CFQ
Assumption: Crossover-point calculation assumes approximately normal group distributions. For heavily skewed or multimodal distributions, interpret with caution.
Note: Primary robustness metric for continuous outcomes. MeCI captures population-level distributional separation independent of α thresholds and sample size, complementing CFQ which measures fragility of the corresponding significance classification. The underlying distance quantity d follows the definition in Heston (2025); the canonical NBF form applies the x/(1+x) bounding map for consistency with the rest of the NBF family (DNB, DTI, ORQ). CI-only implementation note: MeCI cannot be computed from a published mean-difference CI alone, as the crossover-point formula requires separate group means and SDs. In CI-only scenarios, report CFQ alone or reconstruct group-level summary statistics where possible.
Application: Correlation/association studies where the primary result is a correlation coefficient (e.g., Pearson or Spearman r).
Definition: NBF-based robustness metric for correlations.
Formula: DTI = |atanh(r)| / (1 + |atanh(r)|)
Range: 0 to 1.
Interpretation: For DTI, nb = DTI. For example, nb = 0.25 means the correlation is clearly separated from zero.
Neutrality: r = 0
Pairs with: ZFQ
Note: Primary robustness measure for correlation studies; member of the distance-from-neutrality transform family (Part XI).
Application: Multi-group continuous outcome comparisons analysed with one-way ANOVA or equivalent F-tests.
Definition: NBF-compatible robustness metric for multi-group comparisons.
Formula: η² = df_b·F / (df_b·F + df_w)
Range: 0 to 1.
Interpretation: For ANOVAη², nb = η². For example, nb = 0.30 means substantial between-group variation relative to within-group variation.
Neutrality: All group means equal (F = 0).
Pairs with: ANOVA-FQ
Note: Equivalent to the traditional eta-squared effect size; already NBF-compatible. Now paired with ANOVA-FQ to provide complete p–fr–nb triplet for one-way ANOVA designs.
Application: Single-arm proportion vs benchmark analyses (e.g., single-arm response rates or agreement vs benchmark) when only k, n_relevant, and p₀ are available.
Definition: Proportion-NBF measures geometric distance between an observed single-arm proportion and a benchmark proportion p₀. It is the NBF counterpart to BFQ when only k, n_relevant, and p₀ are available.
Formula:
Application: Matched-pair or fixed-margin 2×2 designs (e.g., crossover trials, pre–post paired binary outcomes, or any setting where McNemar’s test is appropriate).
Definition: NBF-style robustness metric for marginal homogeneity; proportion of discordant pairs that would need to switch direction to reach b = c.
Formula: MHQ = |b − c| / (b + c) if b + c > 0 else 0
Range: 0 to 1
Interpretation: nb = MHQ. Example: nb = 0.20 means 20% of discordant pairs must switch direction to reach b = c (neutrality).
Advantages: Matches the McNemar null exactly, and is intuitive for matched/paired designs.
Neutrality: b = c (marginal homogeneity)
Pairs with: McNemar-path fragility (matched-pair designs) = PFI-M (the McNemar-test variant of PFI for matched-pair designs; identical path construction, McNemar χ² in place of Pearson χ²)
Note: Used as the nb metric in fixed-margin/matched-pair modules, paired with McNemar-path fragility so that p, fr, and nb all reference marginal homogeneity. For independent-sample designs, PFI (fragility) and RQ (distance from independence) are the defaults.
Application: Ordinal outcomes analyzed via Wilcoxon-Mann-Whitney test, proportional odds models, or ordinal logistic regression (e.g., modified Rankin Scale, NIHSS, pain scales). Definition: NBF-based robustness metric measuring geometric distance from neutrality (gOR = 1) for ordinal outcomes. Formula: Let gOR = generalized odds ratio (common odds ratio from proportional odds model). Then: ORQ = |ln(gOR)| / (1 + |ln(gOR)|) Range: 0 to 1 Interpretation: nb = ORQ. Example: nb = 0.23 means the ordinal outcome shows moderate separation from neutrality; nb = 0.50+ indicates strong shift toward better outcomes. Neutrality: gOR = 1 (no ordinal shift between groups) Pairs with: OFQ Note: Primary robustness metric for ordinal outcomes. Uses natural log transformation (consistent with DNB for diagnostic odds ratios). Works from published gOR alone—no confidence interval needed for ORQ calculation (though CI is needed for OFQ). Member of the distance-from-neutrality transform family (Part XI).
Application: Time-to-event outcomes analyzed via Cox regression (e.g., overall survival, disease-free survival, cardiovascular mortality). Definition: NBF-based robustness metric measuring geometric distance from neutrality (HR = 1) for survival outcomes. Formula: Let HR = hazard ratio from Cox regression. Then: SRQ = |ln(HR)| / (1 + |ln(HR)|) Range: 0 to 1 Interpretation: nb = SRQ. Example: nb = 0.18 means the hazard ratio shows moderate separation from neutrality; nb = 0.50+ indicates strong reduction (or increase) in hazard. Neutrality: HR = 1 (equal hazard rates between groups; no treatment effect) Pairs with: SFQ Note: Primary robustness metric for survival outcomes. Uses natural log transformation (consistent with DNB for diagnostic odds ratios and ORQ for ordinal outcomes). Works from published HR alone—no confidence interval needed for SRQ calculation (though CI is needed for SFQ). Member of the distance-from-neutrality transform family (Part XI).
Definition: Number of outcome toggles within one arm required to flip statistical significance from significant to nonsignificant. Only defined for baseline statistically significant 2×2 contingency tables. Defined procedurally, not as a minimum over a search space: the index is whatever count the prescribed iterative procedure yields. Toggle rule: Convert non-events to events, one at a time, in the arm with fewer events (recalculating p after each toggle) until p ≥ 0.05; not defined what to do if the event counts are tied. Test: Two-sided Fisher's exact, recalculated at each step regardless of the significance test the original trial reported. Output: Integer count → WalshFQ = WalshFI / N, WalshMFQ = WalshFI / n_mod. Note: Classic metric from Walsh et al. (2014). The original empirical study applied it to RCTs with 1:1 allocation ratios; the index definition itself does not require equal allocation.
Definition: Minimum number of outcome toggles within one arm required to flip the statistical significance classification in either direction (significant→nonsignificant or nonsignificant→significant), using a two-sided Fisher's exact test. Defined for both significant and nonsignificant baseline 2×2 contingency tables. Defined as a true minimum over the allowed toggles, not a procedural count. Toggle rule: Toggle outcomes (event ↔ non-event) in the arm with fewer events; if tied, toggle the smaller arm. Test: Two-sided Fisher's exact. Output: Integer count → FQ = FI / N, MFQ = FI / n_mod. Note: Heston modification of Walsh et al. (2014): bidirectional, defined for nonsignificant baselines, defined as a minimum, with an explicit tie rule.
Definition: Minimum total number of within-arm outcome toggles, distributed across either or both arms, required to flip the significance classification using a two-sided Fisher's exact test. Formally: min(|f₀| + |f₁|), where f₀ and f₁ are the net event-count changes in each arm.
Toggle rule: Toggle outcomes (event ↔ non-event) within an arm, preserving both arm sizes; bidirectional; either or both arms may be modified in the same solution.
Test: Two-sided Fisher's exact (package default; chi-squared and OR/RR/RD-based p-values are settable options).
Output: Integer count → LinFQ = LinFI / N; MFQ does not apply since there is no single modified arm.
Note: Generalizes the classic Walsh et al. (2014) metric as implemented by the fragility R package (Lin & Chu, 2022, frag.study): unlike Walsh FI, it is bidirectional (also defined for nonsignificant baselines, flipping toward significance) and is not restricted to a single pre-specified arm.
Definition: Minimum number of cell-to-cell reallocations required to flip the statistical significance classification in an r×c contingency table, using the same significance test throughout the GFI search. Toggle rule: Any admissible cell movement; algorithm finds the minimal global path. Test: For 2×2 tables, GFI always uses the two-sided Fisher's exact test — the same test used by FI, MFI, and SFI — regardless of expected cell counts and regardless of the significance test the original trial reported. This guarantees test-coherence across the fragility metrics and preserves the structural invariant GFI ≤ MFI ≤ FI (every within-arm toggle is a unit reallocation, so the global minimum can never exceed the constrained minima when all three are scored by the same test). For r×c tables larger than 2×2, use the Fisher–Freeman–Halton exact test when computationally feasible; Pearson's chi-square is an acceptable asymptotic substitute when exact computation is intractable and expected counts are adequate (no expected cell count <1 and at least 80% of expected cell counts ≥5). The selected test remains fixed throughout the search and is not changed as cell counts are reallocated. As a sensitivity analysis only, GFI may additionally be computed under the test the original analysis prespecified; such values must be labeled with the test used and must not be mixed with Fisher-based FI/MFI/SFI values in fragility comparisons, because metrics scored under different tests are not mutually comparable and can violate the GFI ≤ MFI ≤ FI ordering by ±1 or more. Unit: Global Fragility Unit (GFU) = 1/N. Output: Integer count → GFQ = GFI / N. Note: Gold standard for binary and multinomial tables: path-independent, label-invariant, and domain-complete (defined for every table except the small-N floor, where no arrangement of N subjects can reach significance at the chosen α). GFI measures the minimum global perturbation required to reverse the classification produced by the significance test; therefore, the significance criterion must remain test-coherent across the baseline, all candidate tables, and the companion metrics (FI, MFI, SFI) it is reported alongside. For 2×2 tables the criterion is fixed at two-sided Fisher's exact; for larger r×c tables the exact test (Fisher–Freeman–Halton) is preferred, with Pearson's chi-square as the asymptotic fallback for adequately populated tables where exact computation is intractable. Exact-test GFI may be more computationally demanding than chi-square GFI. Because Fisher and chi-square p-values differ near the α boundary, GFI values computed under different tests typically differ by ±1 reallocation and must not be compared directly.
Definition: Minimum total continuous cell movement — measured as L1 distance / 2, i.e., the number of "subjects' worth" of mass reallocated — required to flip the statistical significance classification, in either direction, holding total N fixed and all cells ≥ 0. Cells are treated as continuous quantities, so fractional reallocations are admissible. Move rule: Any cell-to-cell reallocation, in continuous amounts. Computed exactly along all 18 canonical rays (12 single-pair moves, 6 diagonal exchanges) by closed-form polynomial root-finding on the chi-square boundary, then refined by constrained optimization (SLSQP) to capture off-ray minima. Test: Two-sided Pearson chi-square (continuous criterion, no Yates correction). Fisher's exact test is not applicable: it does not allow fractional counts, for the same reason as PFI (§3.6). Unit: Continuous; sub-integer resolution (same interpretive basis as the GFU, but no longer restricted to integer multiples of 1/N in quotient form). Output: Non-negative real number → cGFQ = cGFI / N. Undefined (NaN) only at the small-N floor, where no arrangement of N subjects can reach significance at the chosen α. Note: Continuous relaxation of the GFI. For tables scored by chi-square, cGFI ≤ GFI, since every integer reallocation path is also a continuous one; the gap between them measures how much of the integer index is quantization. Like PFI, cGFI resolves fragility differences among tables that share the same integer GFI (many tables have GFI = 1 yet sit at very different distances from the significance boundary), and like PFI it is a descriptive evidence-quality metric, not an inferential test: the intermediate fractional table need not be realizable. cGFI differs from PFI in move space — PFI is restricted to the single fixed-margin diagonal path, while cGFI minimizes over all continuous reallocation paths, so cGFI ≤ (N/4)·PFI on the common path and is the tighter boundary-distance measure.
Definition: Minimum number of coupled fixed-margin reassignments required to bring a 2×2 table to the reachable point closest to therapeutic neutrality (relative risk = 1, equivalently cross-product difference ad − bc = 0). Because exact neutrality (ad = bc) frequently cannot be realized on the integer lattice, the target is the reachable table minimizing |ad − bc|. Defined for every 2×2 table. Move rule: Within-row transfers only, applied as coupled anti-parallel pairs holding both row and column margins fixed. Forward move (a−1, b+1, c+1, d−1); reverse move (a+1, b−1, c−1, d+1) — equivalently a→b paired with d→c, and b→a paired with c→d. Cross-arm transfers (a→c, b→d) are inadmissible. Bidirectional: whichever coupled direction reaches closest approach in fewer moves is taken. This is the Feinstein/Walter fixed-both-margin move-set (the UFI move-set, Part VII), chosen because Fisher's exact test conditions on both margins, so the margin-preserving move is the perturbation commensurable with the conditional test, and neutrality (ad = bc) is itself a within-margin statement. Formula: Each coupled move changes the cross-product difference by exactly ∓N with both margins and N held fixed — (a−1)(d−1) − (b+1)(c+1) = (ad − bc) − N — so reachable cross-product differences are spaced N apart and NDI = round(|ad − bc| / N) = round(N·RQ/4), clamped to the reachability window. Exact neutrality is attainable only when N divides (ad − bc). Reachability: The coupled move is bounded to x ∈ [−min(a, d), +min(b, c)]. Extreme or lopsided margins yield a narrow window, and NDI then reflects how little the table can move within Fisher's conditioning set. Test dependence: None. NDI depends only on ad − bc, a function of the observed cell counts; it invokes no significance test. This distinguishes it from the fragility counts (FI, GFI), whose target is a p-value threshold. NDI is a robustness metric expressed as an integer count. Domain: Always defined. Unlike the fixed-margin UFI — undefined when no reachable table flips significance — NDI's target is a minimization that always has a solution. NDI = 0 is a valid result: the observed table is already as close to neutrality as the fixed-margin moves permit. Unit: One coupled move = 1 unit, matching the published UFI. One coupled move relocates two patients (one per arm), so the fixed-margin indices (UFI, NDI) are on a per-coupled-move unit whereas single-transfer indices (GFI) are on a per-patient unit. The difference is by design, because the move-sets differ. Output: Integer count, range 0 to N/4 (since |ad − bc| ≤ N²/4). NDI is the count-based robustness metric; RQ (§4.1) is its normalized decimal partner, so NDI:RQ parallels the GFI:GFQ Index:Quotient pairing. Note: NDI = round(N·RQ/4) is an exact algebraic identity, not an empirical approximation — it follows from RQ = |ad − bc|/(N²/4) and the ∓N step size. The only departures from exactness are integer rounding (coarse near neutrality, where the lattice spacing N is large relative to a small |ad − bc|) and clamping at extreme margins where the reachability window binds. RQ is therefore not NDI/N. Note: Robustness counterpart to the fixed-margin unit fragility index (UFI, Part VII): UFI counts coupled fixed-margin moves to the significance boundary (p = 0.05); NDI counts the same coupled moves to the neutrality boundary (RR = 1). Together they instantiate both boundaries of the framework with one move-set. NDI is model-free (computed from the 2×2 summary counts alone, no distributional assumptions) and measures distance-to-neutrality in patient units, which p-values do not measure; it is not a transformation of p.
Definition: Minimum number of "success" toggles required to switch the diagnostic benchmark classification between "below benchmark" and "not below benchmark" (one-sided exact binomial).
Toggle rule:
• Sensitivity: TP ↔ FN
• Specificity: TN ↔ FP
• PPV: TP ↔ FP
• NPV: TN ↔ FN
• Accuracy: any TP/TN/FP/FN toggle changing success proportion
Test: One-sided exact binomial test vs benchmark p₀.
Output: Integer count → DFQ = DFI / n_relevant.
Note: Underlying count for DFQ; always paired with DNB for diagnostic accuracy studies.
Definition: Minimum number of success/failure toggles in a single-arm binomial experiment required to switch the benchmark classification between "meets benchmark" and "does not meet benchmark" under a one-sided exact binomial test.
Toggle rule: Toggle individual outcomes (success ↔ failure) in the single arm until the one-sided exact binomial decision crosses the α = 0.05 boundary at the specified benchmark p₀.
Test: One-sided exact binomial test vs benchmark p₀.
Output: Integer count → BFQ = BFI / n_relevant (with n_relevant = n).
Note: Underlying count for BFQ in single-arm proportion vs benchmark analyses; used together with Proportion-NBF as the robustness partner.
Definition: SE-unit distance between the observed Welch t-statistic and the α = 0.05 significance boundary.
Formula: CFS = ||T| – t*|.
Output: Raw distance → CFQ = CFS / (1 + CFS).
Note: Continuous analogue of FI/SFI/GFI.
Welch model (two independent continuous groups):
v₁ = s₁² / n₁
v₂ = s₂² / n₂
CFU = SE_diff = √(v₁ + v₂)
Purpose: One SE-shift in the estimated mean difference under Welch.
CFS = ||T| − t*|
Defines the number of CFUs needed to reach the p = 0.05 boundary (see §3.7).
Base metric for CFQ.
Definition: Raw geometric distance from independence in a multinomial table, defined as the average absolute difference between observed and expected cell counts.
Formula: RRI = (1/k) Σ|O − E|, where k is the number of cells (k = r×c), O are observed counts, and E are expected counts under independence.
Output: Distance value → RQ = k·RRI / [2N(m − 1)/m], m = min(r, c); since k·RRI = Σ|O − E|, this equals the general RQ of §4.1. The shortcut RQ = RRI / (N/k) = Σ|O − E| / N holds only when m = 2 (any 2×c or r×2 table, including 2×2), where 2N(m − 1)/m reduces to N; for min(r, c) ≥ 3 it does not, and the general form applies.
Note: Parent metric for RQ.
Definition: Let N denote the total sample size and α the significance threshold (default 0.05). SFM is the smallest scaling factor k > 1 such that multiplying N by k flips a nonsignificant result to significant, or dividing N by k flips a significant result to nonsignificant.
Purpose: Sample-size sensitivity of significance classification.
Output: k > 1.
Interpretation: Values near 1 indicate fragile significance status; larger values indicate greater stability.
Note: Exploratory only; superseded by nb (robustness). Renamed from “RI” in previous versions to correctly classify as a fragility metric.
Definition: Fragility count in standardized binomial fragility units (BFUs) where BFU = 1/n_large (n_large is the number of subjects in the larger arm).
Toggle rule: Toggle outcomes in the arm with more subjects (or if tied, the fewer-events arm) until significance reverses.
Test: Two-sided Fisher's exact.
Output: Integer count → MFQ_SFI = SFI / n_large.
Note: Valid but redundant; MFQ is preferred.
Definition: Fixed-margin fragility framework for matched or hypergeometric designs.
Feinstein (Unit Size):
For a 2×2 fixed-margin (hypergeometric) table, the minimal admissible toggle size is
f = N / (n₁ n₂),
defining the smallest allowable perturbation under fixed row/column totals.
Walter (Toggle Count):
Given unit size f, Walter defines UFI as the minimum number k of these fixed-margin unit shifts required to reverse significance.
Output:
Note: Both strictly fixed-margin constructs. Conceptual precursors to PFI. Modern analyses use PFI for fixed margins and MFQ/GFQ otherwise.
Note: The Walter coupled fixed-margin move-set is reused by NDI (Part V), which counts the same moves to the neutrality boundary (RR = 1) rather than to the significance boundary (p = 0.05). UFI and NDI therefore instantiate both boundaries of the framework with a single move-set.
Definition: Variant allowing within-arm toggles in either arm, taking the minimum count required to reverse significance.
Note: Adds no practical value beyond FI → MFQ. Retained only for historical completeness.
Fragility Metrics (0–1)
Robustness Metrics (0–1)
Interpretation depends on the claim being made:
| Metric | Desirable | Problematic |
|---|---|---|
| Fragility | High quotient (stable p-value) | Low quotient (unstable) |
| Robustness | High NBF (far from neutral) | Low NBF (near neutral) |
Best case: High fragility quotient + High robustness
Worst case: Low fragility quotient + Low robustness
| Metric | Desirable | Problematic |
|---|---|---|
| Fragility | High quotient (stable p-value) | Low quotient (unstable) |
| Robustness | Low NBF (near neutral) | High NBF (far from neutral) |
Best case: High fragility quotient + Low robustness
Worst case: Low fragility quotient + High robustness
To understand complete statistical evidence, consider an exam where:
A score of 61% means you barely passed—a few unlucky questions could have failed you. A score of 85% means you passed convincingly. Similarly, fragility measures how close your p-value sits to the 0.05 decision boundary.
Mastery measures something different: your true underlying knowledge. A student with 30% mastery is barely above random guessing. A student with 85% mastery genuinely knows the material. Robustness similarly measures how far your observed result sits from "no effect"—the neutrality boundary.
These can diverge. A knowledgeable student can score poorly (bad luck on questions asked). A weak student can score well (lucky guesses or easy test). The same applies to statistical evidence.
| Score | Mastery | Triplet | Interpretation |
|---|---|---|---|
| 61% | 30% | p-sig, fr-fragile, nb-weak (Pattern 1,1,0) | Barely passed with minimal knowledge. The pass is legitimate but mastery is negligible. Thin evidence—potentially real mastery, but trivial competence. Do not certify for practice. |
| 61% | 55% | p-sig, fr-fragile, nb-moderate | Barely passed, knows something. Inconclusive—may or may not replicate. |
| 61% | 85% | p-sig, fr-fragile, nb-strong | Barely passed despite strong knowledge. Unlucky draw of questions. Underpowered true positive—effect is real but study was too small to detect it reliably. |
| 85% | 30% | p-sig, fr-stable, nb-weak | Passed easily but knows little. Test was too easy. Statistically robust but trivial mastery—overpowered detection of negligible competence. |
| 85% | 55% | p-sig, fr-stable, nb-moderate | Solid pass, moderate knowledge. Good evidence of real, modest mastery. |
| 85% | 85% | p-sig, fr-stable, nb-strong | Clear pass, clear mastery. Compelling evidence. This is the goal. |
| Score | Mastery | Triplet | Interpretation |
|---|---|---|---|
| 59% | 30% | p-nonsig, fr-fragile, nb-weak | Barely failed, doesn't know much. Probably true negative, but verdict is unstable. |
| 59% | 55% | p-nonsig, fr-fragile, nb-moderate | Barely failed, knows something. Inconclusive—needs more data. |
| 59% | 85% | p-nonsig, fr-fragile, nb-strong | Barely failed despite strong knowledge. Bad luck on questions. Likely false negative—effect exists but was missed. |
| 45% | 30% | p-nonsig, fr-stable, nb-weak | Clearly failed, doesn't know much. Strong evidence of no meaningful competence. True negative. |
| 45% | 55% | p-nonsig, fr-stable, nb-moderate | Clearly failed but has some knowledge. Underpowered—a real effect may have been missed. |
| 45% | 85% | p-nonsig, fr-stable, nb-strong | Clearly failed despite clearly knowing material. Severely underpowered—test design was fundamentally inadequate to assess this student. |
In exams: limited questions, bad luck, wrong format, test too easy or too hard.
In studies: small samples, high variance, inadequate power, overpowered detection of trivial effects.
The p-value (pass/fail) captures only part of the picture. Complete statistical evidence requires all three dimensions.
Thresholds are recommendations and still require empirical validation and should be treated as provisional.
Numeric fragility cutoffs (for example, an MFQ "fragile vs stable" threshold) are under validation and are omitted here pending publication. Interpret fr qualitatively — lower = more fragile, higher = more stable — within each design family.
RQ Percentiles (1M mixed simulated trials)
| 1% | 5% | 10% | 25% | 33% | 50% | 67% | 75% | 90% | 95% | 99% |
|---|---|---|---|---|---|---|---|---|---|---|
| 0.0018 | 0.0092 | 0.0188 | 0.0525 | 0.0746 | 0.1358 | 0.2274 | 0.2875 | 0.4322 | 0.4942 | 0.5825 |
Proposed Empirical Cutoffs:
Weak robustness (near neutrality): RQ < 0.075
Moderate robustness: RQ 0.075 - 0.227
Strong robustness (far from null): RQ > 0.227
| Range | Distance from Neutrality |
|---|---|
| < 0.075 | Close to neutrality / Weak Robustness |
| 0.075-0.227 | Moderate separation / Moderate Robustness |
| ≥ 0.227 | Far from neutrality / Strong Robustness |
The following strength-of-evidence tables are calibrated for 2×2 binary trials using MFQ. For other designs, analogous tables use the native fragility quotient for that design, interpreted within its family.
When p ≤ 0.05 (statistically significant)
| Fragility | Robustness | Interpretation | Diagnosis | Action |
|---|---|---|---|---|
| Stable | Strong (RQ high) | Significant, robust, large effect | Reliable detection of substantial effect. | ✅ Trust for clinical use* |
| Stable | Moderate | Stable significance with moderate effect | Reliable detection of modest effect. Clinical significance depends on absolute benefit and baseline risk. | ✅ Consider if effect size clears clinical threshold |
| Stable | Weak (RQ low) | Stable significance but near neutrality | Trivial effect reliably detected. Statistically significant ≠ clinically meaningful. | ⚠️ Caution - effect too small |
| Fragile | Strong | Fragile yet far from neutrality | Significant but unstable. Effect appears real but easily overturned. | ⚠️ Caution - verify in larger sample |
| Fragile | Moderate | Classic fragile result | Unstable effect of uncertain magnitude. Could be real, could be noise. | ⚠️ Replicate before use |
| Fragile | Weak | Pattern (1,1,0) | Thin evidence. Effect uncertain. Significant p-value provides false confidence. | ⛔ Reject for clinical use |
When p > 0.05 (statistically nonsignificant)
| Fragility | Robustness | Interpretation | Diagnosis | Action |
|---|---|---|---|---|
| Stable | Weak (RQ low) | Stable nonsignificance near neutrality | Strong evidence of no effect. Reliable null result. Effect truly absent or negligible. | ✅ Trust the null |
| Stable | Moderate | Stable nonsignificance with moderate effect | Possible effect not reaching significance. Directionally consistent but underpowered. | ⚠️ Consider replication if clinically important |
| Stable | Strong (RQ high) | Stable nonsignificance yet far from neutrality | Severely underpowered. Effect clearly exists but remains nonsignificant. Design inadequate to detect real effect. | ⛔ Underpowered - need larger trial |
| Fragile | Weak | Fragile nonsignificance near neutrality | Likely true negative but unstable. Probably no effect, though classification fragile. Leans toward null. | ✅ Likely null (low confidence) |
| Fragile | Moderate | Fragile nonsignificance with moderate effect | Inconclusive. Cannot distinguish "no effect" from "missed effect." Borderline case. | ⚠️ Inconclusive - need more data |
| Fragile | Strong | Fragile nonsignificance yet far from neutrality | Likely false negative. Effect exists but wasn't detected. Just missed significance threshold. | ⛔ False negative - increase power |
Key:
Critical Note: These classifications assess statistical robustness and proximity to neutrality. Always evaluate absolute effect sizes (NNT, risk difference, absolute risk reduction) and baseline risk before clinical application. A statistically robust finding with clinically negligible absolute magnitude still warrants caution.
Quick Reference: What Went Wrong?
| Pattern | Problem | Solution |
|---|---|---|
| p-sig, fr-stable, nb-weak | Trivial effect reliably detected | Question clinical relevance; report absolute effect size |
| p-sig, fr-fragile, nb-strong | Underpowered for real effect | Replicate with adequate power |
| p-sig, fr-fragile, nb-weak (Pattern 1,1,0) | Evidence consistent with a trivial effect | Treat as clinically negligible unless replicated with stronger nb. |
| p-nonsig, fr-stable, nb-strong | Severely underpowered despite real effect | Redesign with proper power calculation |
| p-nonsig, fr-fragile, nb-strong | Underpowered, missed real effect | Replicate with adequate power |
| p-nonsig, fr-stable, nb-weak | Nothing wrong—correct null finding | Trust the null |
| p-sig, fr-fragile, nb-moderate | Classic borderline result—uncertain magnitude | Replicate before clinical implementation |
| p-nonsig, fr-fragile, nb-moderate | Inconclusive—cannot distinguish null from missed effect | Need more data to resolve |
Note: This table highlights problematic patterns requiring action. Omitted patterns (p-sig with stable+strong, stable+moderate, or fragile+moderate) represent acceptable findings when absolute effects are clinically meaningful.
FQ = FI / N
MFQ = FI / n_mod
GFQ = GFI / N
DFQ = DFI / n_relevant
BFQ = BFI / n_relevant (with n_relevant = n in single-arm designs)
CFQ = CFS / (1 + CFS)
PFI = 4·|x| / N
ANOVA-FQ = ANOVA-FS / (1 + ANOVA-FS)
ZFQ = D / (1 + D)
OFQ = | |z_WMW| − 1.96| / (1 + | |z_WMW| − 1.96|)
SFQ = | |z_HR| − 1.96| / (1 + | |z_HR| − 1.96|)
NDI = round(N·RQ/4) = round(|ad − bc| / N) (2×2; clamped to the reachability window)
GFI ≤ FI (always)
GFI ≤ SFI (always)
0 ≤ NDI ≤ N/4 (always, since |ad − bc| ≤ N²/4)
All quotients: in [0,1]
All NBF metrics: in [0,1]
□ All three dimensions reported (significance, fragility + robustness)
□ Primary metrics used (quotients, not just counts)
□ Effect size with 95% CI included
□ Exact p-value reported
□ Interpretation matches the stated claim
□ Prevalence normalized for PPV/NPV when appropriate
□ For continuous outcomes, CFQ + MeCI reported when possible
The modern statistical evidence framework consists of three complementary dimensions, all scaled 0–1:
PROBABILITY (p): the p-value from the design's significance test — compatibility of the observed data with the null (no effect). Lower p = stronger evidence against no effect. Determines a study's significance classification: p < 0.05 = significant, p ≥ 0.05 = not significant.
FRAGILITY (quotient-based): What proportion must change to flip p?
ROBUSTNESS (NBF-based): How far from neutrality?
Interpretation depends on the claim:
The robustness and fragility dimensions at the meta-analytic level are not separate metrics. For binary-outcome meta-analyses they are operationalized as N-weighted pooled scalars over the per-study metrics already defined in this document: wsRQ (weighted signed RQ) for the robustness dimension and wGFQ (weighted global fragility quotient) for the fragility dimension. All inputs are computed from the aggregate event/non-event counts already published for each component study; no patient-level data, simulation, or distributional assumptions are required.
Definition: N-proportional weighting. w_i = N_i / ΣN_j, where N_i is the sample size of component study i and the sum runs over all contributing studies (Σw_i = 1). Constraint: The weight MUST be N-proportional. It is neither inverse-variance nor Mantel–Haenszel. N-weighting preserves RQ's model-free, estimator-independent property (see advantages note below); inverse-variance or Mantel–Haenszel weighting would forfeit it by reintroducing dependence on the between-study variance estimator.
Application: Meta-analytic robustness for binary outcomes; the signed, pooled distance from therapeutic neutrality across component studies. Formula: wsRQ = Σ(sRQ_i · w_i), where sRQ_i = 4(ad − bc)/N² for study i (§4.1.1) and w_i is the N-proportional weight. Range: −1 to +1. Direction: Inherits the sRQ orientation convention (§4.1.1); the same table orientation must be applied to every component study before pooling. Sign indicates the net favored direction across the synthesis. Neutrality: wsRQ = 0 (pooled independence). Pairs with: wGFQ.
Application: Meta-analytic fragility for binary outcomes; the pooled proportion-to-flip across component studies. Formula: wGFQ = Σ(GFQ_i · w_i), where GFQ_i = GFI_i / N_i (§3.3) and GFI_i is computed per the Dijkstra two-phase calculator (Fragility Metrics Toolkit, Zenodo 10.5281/zenodo.17254763), and w_i is the N-proportional weight. Range: 0 to 1. Neutrality: not applicable (fragility metric; higher = more stable classification). Pairs with: wsRQ.
These pooled weighted scalars (wsRQ, wGFQ) are the primary meta-analytic operationalization. Pooling is performed over the per-study metrics, not by summing events and non-events into a single pooled 2×2 table.
Worked regression values (Zuin et al. PE-thrombolysis reanalysis; 10 trials for all-cause mortality, 9 trials for major bleeding): mortality wsRQ = −0.011457, wGFQ = 0.018857; bleeding wsRQ = +0.054340, wGFQ = 0.022637. Both outcomes pool to fragile and weakly robust.
A practical advantage: sRQ/RQ and GFQ are computed independently of the P-value and of the between-study variance estimator, so the pooled scalars do not shift when a meta-analytic FI shifts under different model choices (fixed vs. random effects, REML vs. DerSimonian-Laird, HKSJ CI adjustment). The N-proportional weight operator preserves this estimator-independence; inverse-variance or Mantel–Haenszel weighting would not. See Heston TF, Reverse Fragility in Cochrane Meta-Analyses with P Values 0.05 to 0.20 Requires a Robustness Dimension, Internet Med J. 2026;1:e19741629 (doi:10.5281/zenodo.19741629).
A prior operationalization (v11.1.2) sorted each nonsignificant meta-analysis into one of six interpretive cells defined by fragility (fragile vs. stable) × distance from independence (weak / moderate / strong), aggregating per-study triplets without pooling to a scalar. This categorical sort is superseded by the pooled weighted scalars (wsRQ, wGFQ) above and is retired. It is retained here only as a record of the supersession, not as a recommended method. Note: the published reverse-fragility commentary (Internet Med J 2026;1:e19741629, doi:10.5281/zenodo.19741629) used the prior categorical framing, so canon now diverges from that piece.
Note: "nb_meta" is not a defined metric in this framework. Earlier internal tracking notes used "nb_meta" as informal shorthand for the meta-analytic robustness dimension; the correct operationalization for binary outcomes is the pooled weighted scalar wsRQ (paired with wGFQ), as specified above.
The framework's metrics are organized into families to give the standing methods-note series a principled per-family structure. Family boundaries:
FQ family (native fragility quotients): GFQ, MFQ, CFQ, SFQ, BFQ
RQ family (NBF robustness, independence/risk axis): RQ, sRQ, SRQ, MHQ
Meta-analytic variants (N-weighted pooled scalars + weight operator): wsRQ, wGFQ, weight operator w_i
Index forms (raw counts/distances): GFI, FI, MFI, SFI, UFI, PFI, NDI (NDI is the sole robustness-target member — it counts moves to neutrality rather than to significance)
Distance-to-critical-value family (fragility, form δ/(1+δ) with δ = |test statistic − α-critical value|): CFQ, ANOVA-FQ, ZFQ, OFQ, SFQ — one construct instantiated per design; each keeps its own statistic and critical value.
Distance-from-neutrality transform family (robustness, form |g(θ)|/(1+|g(θ)|) with g a variance-stabilizing transform of the effect estimate θ): DTI (g = atanh, θ = r), ORQ (g = ln, θ = gOR), SRQ (g = ln, θ = HR), DNB (g = ln, θ = DOR) — one construct instantiated per design.
These groupings are organizational for the methods-note series and do not override any individual metric's canonical definition or NBF pairing stated elsewhere in this document.
Every methods note in the standing series is uniform and Scholar-optimized. Required structure:
Title: contains the metric or family name verbatim.
Body sections (in order):
Standard declarations block: funding, conflict of interest, and AI-use declarations.
References: APA7 format; the hub citation appears within the first three references.
Ahmed W, Fowler RA, McCredie VA. Does sample size matter when interpreting the fragility index? Crit Care Med. 2016;44(11):e1142–3.
Defines the classic Fragility Quotient (FQ = FI/N) and highlights the dependence of FI on sample size.
Baer BR, Gaudino M, Charlson M, Fremes SE, Wells MT. Fragility indices for only sufficiently likely modifications. Proc Natl Acad Sci USA. 2021;118(49):e2105254118.
Because their methods require model assumptions, probability weighting, or reconstructed data, the resulting fragility measures stop being properties of the evidence and become properties of the chosen model. The purpose of this framework is to preserve fragility and robustness as direct, model-free functions of the observed data and exact tests. Anything that introduces subject-level probabilities, covariate structures, or simulated counterfactuals breaks that principle.
Caldwell JME, Youssefzadeh K, Limpisvasti O. A method for calculating the fragility index of continuous outcomes. J Clin Epidemiol. 2021;136:20–25.
Introduces the Continuous Fragility Index (CFI), which perturbs raw data or generates pseudo–individual observations from summary statistics under distributional assumptions. This reconstruction step makes the fragility measure depend on the modeling choices rather than the observed evidence. In contrast, the CFQ/CFS framework uses only published summary statistics and the exact Welch test geometry, producing a unique, model-free fragility value. CFQ is therefore preferred because it is reproducible, assumption-free, and aligned with the binary and multinomial fragility definitions.
Heston TF. Adjusting fragility metrics for unequal trial randomizations. Autoimmun Rev. 2025;24(12):103935.
Demonstrates that classic fragility measures can misrepresent stability when treatment arms are imbalanced and formalizes the allocation-corrected adjustment that underlies MFQ. Provides the empirical and mathematical justification for normalizing fragility to the arm actually subjected to toggling, resolving the asymmetry and mis-scaling inherent in FQ for unequal randomizations.
Heston TF. Fragility Metrics Toolkit Zenodo. 2025;17254763.
Open-source reference implementation containing FI, FQ, MFQ, GFI, GFQ, PFI, UFI, and RQ. Establishes computational standards for the core model-free fragility and robustness metrics currently available. Additional metrics (DFI/DFQ, CFS/CFQ, DNB, MeCI, DTI, ANOVA-FQ, ZFQ, ANOVAη²) were originally developed outside the core toolkit and are now integrated into this reference; software implementations will follow in subsequent toolkit releases.
Heston TF. Meaningful Change Index: A P-Value Independent Metric for Assessing Robustness and Fragility in Continuous Outcomes. SSRN. 2025;5535978.
Defines the MeCI robustness metric for continuous outcomes as the minimum distance from group means to their distributional crossover point, normalized by combined standard deviations. Establishes MeCI as p-value independent and sample-size independent, and frames it as the robustness complement to CFQ for continuous-outcome trials. The canonical NBF implementation (this document, §4.3) applies an x/(1+x) bounding map to the primary-paper distance quantity for family consistency with RQ (binary) and DNB (diagnostic).
Heston TF. The Global Fragility Index: A Path-Independent Measure of Statistical Fragility. SSRN. 2025;5709162.
Defines the GFI framework for multinomial tables and proves path-independence of the global cell-move distance to the significance boundary. Basis for GFQ and the GFU unit.
Heston TF. The Modified-Arm Fragility Quotient: An Improved Metric for Assessing Robustness in Clinical Trials. SSRN. 2025;5425334.
Establishes MFQ as the allocation-fair fragility quotient for 2×2 trials, showing that FI should be normalized to the arm actually subjected to toggling. This resolves the long-standing imbalance and label-dependence of the classic FQ.
Heston TF. The Neutrality Boundary Framework: Quantifying Statistical Robustness Geometrically. arXiv. 2025;2511.00982.
Introduces the NBF formulation nb = |T − T₀|/(|T − T₀| + S), establishing a unified 0–1 robustness scale for binary, diagnostic, correlation, and multi-group analyses. Provides the mathematical basis for RQ, DNB, DTI, Proportion-NBF, and ANOVAη². MeCI, the NBF robustness metric for continuous outcomes, uses a distributional crossover-point construction rather than the |T−T₀|/(|T−T₀|+S) template but shares the unified 0–1 family scaling via the x/(1+x) bounding map.
Heston TF. Redefining significance: robustness and percent fragility indices in biomedical research. Stats. 2024;7(2):537–48.
Develops PFI for fixed-margin designs and motivates the joint use of fragility (fr) and robustness (nb) as complementary evidence dimensions, anticipating the unified fragility–robustness system formalized in v9.0.
Khan MS, Fonarow GC, Friede T, Lateef N, Khan SU, Anker SD, et al. Application of the reverse fragility index to statistically nonsignificant randomized clinical trial results. JAMA Netw Open. 2020;3(8):e2012469.
Reverse FI extends the classic FI toggling logic to nonsignificant results. This extension does not require a separate fragility construct because fragility can be defined uniformly as the minimal perturbation required to cross the significance boundary in either direction. The Heston FI formalizes this bidirectional definition while retaining a single-arm toggle rule for both significant and nonsignificant baseline tables. Creating a separate “reverse” metric therefore duplicates the underlying mechanism without adding theoretical clarity. The unified fragility framework (MFQ/GFQ/CFQ/DFQ) further removes the need for a significant-versus-nonsignificant distinction by treating fragility as classification stability regardless of which side of the significance boundary the observed result occupies.
Lin L, Chu H. Assessing and visualizing fragility of clinical results with binary outcomes in R using the fragility package. PLoS ONE. 2022;17(6):e0268754.
Implements a modified FI in which both arms are toggled independently, rather than restricting toggles to the fewer-events (or smaller) arm as defined in the original FI procedure. This alters the data-generating assumptions behind FI and breaks comparability across studies. The model-free framework in this reference retains the classic FI toggle rule because it preserves invariance, reproducibility, and direct interpretability; MFQ is built intentionally on that stable foundation rather than on an alternative toggling heuristic.
Walsh M, Srinathan SK, McAuley DF, Mrkobrada M, Levine O, Ribic C, et al. The statistical significance of randomized controlled trial results is frequently fragile: a case for a Fragility Index. J Clin Epidemiol. 2014;67(6):622–8.
Defines the classic FI and the canonical toggle rule on which MFQ is based.
Version 13.3.0 (August 16, 2026)
Version 13.2.1 (August 15, 2026)
Version 13.1.0 (August 13, 2026)
Version 13.0.4 (August 10, 2026)
Version 13.0.3 (July 20, 2026)
Version 13.0.2 (July 20, 2026)
Version 13.0.1 (July 19, 2026)
Version 12.0.0 (July 13, 2026)
Version 11.2.0 (June 25, 2026)
Version 11.1.3 (June 21, 2026)
Version 11.1.2 (June 14, 2026)
Version 11.0.0 (December 17, 2025)
Version 10.3.6 (December 15, 2025)
Version 10.3.5 (December 4, 2025)
Version 10.3.4 (December 3, 2025)
Version 10.3.3 (November 29, 2025)
Version 10.3.2 (November 28, 2025)
Version 10.3.1 (November 27, 2025)
Version 10.3.0 (November 25, 2025)
Version 10.0–10.2 (October–November 2025)
Version 9.x (September–October 2025)
Version 1.0–8.x (2024–2025)
License: CC-BY-4.0
© 2025 Thomas F. Heston
Preferred citation: Heston TF. Fragility Metrics Toolkit. Zenodo. 2025. https://doi.org/10.5281/zenodo.17254763