ReproAI Reproducibility Check — The Fragility Principle (RSOS)

Header

Field Value
Engine Qwen/DeepSeek-V4-Flash Statistics auditor
Rules ReproAI/REPRO_STANDARDS.md
Audit date 2026-09-16
Package FragilityPrinciple v1.0 — Zenodo doi:10.5281/zenodo.21834953 (MSID RSOS-261834)
Mode Live (ran author code end-to-end) + static + independent reimplementation
Host env win32 · Python 3.8.10 (numpy 1.24.4, scipy 1.10.1, pandas 2.0.3, matplotlib 3.7.5) · R 4.6.1 (Rscript at C:\Program Files\R\R-4.6.1\bin)
Host toolchain R and Python are both installed and used on this system. The author code is Python and was re-run live; the independent second-language recompute was done in base R (no author modules). Other statistical packages (Stata, SPSS, MATLAB) are not installed here; in a companion recheck they would need to be translated to R or Python (both available) rather than run natively.
Version note Package declares Python≥3.10 / numpy≥1.26. Re-ran on older ABI-compatible stack with fixed seeds; all values reproduce at reported precision (only ~1e-15 float/BLAS drift).
Scope fragility-principle-v1.0 (code 15 files, results 14 files). No git history (Zenodo snapshot).

Bottom line

Severity Count
P0 blocker 0
P1 high 0
P2 medium 5
P3 low 3

Verdict. This is a genuinely reproducible package of unusually high quality. Every
table (1–4), the abstract statistics, the §3.2 critical ratios, and both §3.3 EEG
confirmatory arms regenerate exactly at reported precision from the shipped seeds,
and the confirmatory classification is independently confirmed from the shipped raw
file with hand-written logic. There are no wrong numbers: no P0/P1. The findings are
all traceability/provenance and scope issues — the strongest being that three
quantitative prose claims (6000-family geometry check, null VR≈0.66 calibration, the
Metric-T "R exceeds |T| threefold" claim) are not backed by any shipped script or
data
, and the promised Metric-T DC(M)/T/R values are never reported in §3.4.

Independent reimplementation table

Script (mine / shipped) Re-implements Manuscript exhibit Status
bia_sobol.py (7) Two-layer functional ANOVA + VR Table 1, Fig 3–4 ✅ exact
bia_sobol3.py (7) Three-layer functional ANOVA, PS scale Table 2, ESM S3 ✅ exact
bia_confirmatory.py (7) Hedges'-g response fit + cluster bootstrap Table 3, §3.2 ✅ exact
okimo_table6.py (7) Empirical 672-spec grid, FANOVA, VR Table 4 ✅ values match; ⚠ schema/provenance drift
eeg_confirmatory.py (8, 0x7) EEG reconstruction arms §3.3 ✅ exact narrative
fragile_geometry.py (7) Geometry + weighted-Φ range §2.4 ✅ kappa=2.5 balance confirmed
sobol3.py --self (7) ANOVA exactness self-tests §2.5 ✅ pass
independent_recompute.py (mine, numpy-only) Re-derive Table 3 coefficients + r* from raw CSV §3.2 ✅ agree (<3e-5; critical ratios exact)
independent_recompute.R (mine, base R) Re-derive Table 3 coefficients, 2000-resample cluster bootstrap CIs, verdicts, classification + r* from raw CSV §3.2/Table 3 ✅ 7/7 classifications OK; all verdicts & r* exact (diff=0.0000)
delay_transition.py (8) ESM §S1 delay sweep ESM S1 ⚠ analysis not runnable (raw not shipped); derived ✓
make_figures.py (7) Figures 1–4 Figs 1–4 ✅ all four build from stored CSVs

Legend: ✅ identical at reported precision · ≈ close/method-explicable · ⚠ discrepant · ➖ not independently checkable.

Numerical reproducibility (headline exhibits)

Table 1 (bia_sobol_results.csv) — all ✅

Manuscript vs shipped/recomputed (medians): null S_E 0.570/0.570, ST_E 0.740/0.740, VR 0.046/0.046; effect-only 0.744/0.069; moderate 0.882/0.364; strong S_E 0.702/0.702, ST_E 0.908/0.908, VR 0.984/0.984. Percentiles agree. Abstract "0.74→0.91", "VR 0.046→0.984" ✅.

Table 2 (bia_sobol3.csv) — all ✅

null ST_E 0.598/0.598, ST_I 0.138/0.138; strong ST_E 0.886/0.886, ST_I 0.008/0.008. §3.1 "inference total effect 0.138→0.008; estimator 0.598→0.886" ✅. Per-replicate component sums = 1.0 (asserted) ✅.

Table 3 (bia_confirmatory.csv) — all ✅

Xc a 0.137[0.128,0.145] nuisance-sensitive; PhA 0.159 nuisance-robust; Xc_H 0.130[0.122,0.138] nuisance-sensitive; R_H/Z_H/BIVA non-target; XcZ nuisance-robust. §3.2 critical ratios r_BMI Xc +1.002[0.922,1.091], Xc_H +0.981[0.902,1.068]; r_H Xc_H +1.698[1.482,1.952] — all regenerate exactly; independent numpy recompute from raw CSV agrees (<3e-5 for coefficients, 0.0000 for critical ratios). Independently confirmed in base R (R 4.6.1): a from-scratch read of bia_confirmatory_raw.csv + 2000-resample cluster bootstrap (seed 20260806) reproduces all seven index classifications (7/7), every verdict, and every critical ratio to 4 decimals (diff = 0.0000). ✅

Table 4 (okimo_table6.csv) — ✅ values, ⚠ provenance

T7-Pz S_E 0.410, S_EP 0.279, ST_I 0.158, VR 0.124, p_VR 0.695, D 0, Φ undefined, Φ⁰ 0 = shipped. AF3-T7 similarly ✅. But current okimo_table6.py no longer produces the p_VR column (it writes an 18-column schema); the p_VR values trace to okimo_fragility.py analyse() (seed 20260807). No single shipped script regenerates the full shipped row set.

Abstract & §3.3

"none of 672 specifications reached significance" — min adjusted p = 0.0603 (T7-Pz) and 0.0503 (AF3-T7), both > 0.05, D=0 ✅. §3.3 common-source arm: six indeterminate + ciPLV equivalent target loadings, all nuisance material ✅; cross-mix arm: 6 material target loadings, empty R_robust ✅ ("no common consensus interval").

Sign-extraction verification (672-spec grid)

The diagnostics rely on per-specification sign of the directed effect. A common
bug class is taking the sign of a two-sided statistic or of |T|; this package does
not do that.

Code path (fragility_sim.py:237-299):
- sign = int(np.sign(g)) where g = (x_B.mean() − x_A.mean())/s_p — the directed
Hedges' g (group B = AB, group A = CB), not |g|, not the p-value.
- Mann–Whitney decision_sign uses the rank direction sign( U/(n_A n_B) − 0.5 )
(probability of superiority), not the (two-sided) p.
- Φ⁰ (phi_zero, line 313) counts r.sign — the directed-effect sign — for all
specs; Φ (fragility_indices) counts decision_sign among significant specs.

Empirical distribution of signed effects (sign_check.py, confirmed against an
independent per-cell recomputation):

Pair g range positive (AB>CB) negative zero mean g Φ⁰
T7-Pz +0.050 … +0.954 672 0 0 +0.4365 0
AF3-T7 +0.031 … +0.785 672 0 0 +0.4207 0

Every one of the 672 directed effects is positive; none is negative or zero.
- r.sign == sign(g) for all 672 specs (True) — Φ⁰ uses the correct directed sign.
- r.decision_sign == independent recomputation (True) — Φ's sign is also correct.
- 30 (T7-Pz) / 22 (AF3-T7) Mann–Whitney specs have a rank-based direction that
opposes the mean direction (decision_sign ≠ sign(g)) — confined to MW rows. This
is intended (rank vs mean functionals) and does not affect Φ⁰ (Φ is undefined because
D=0).

Why Φ⁰ = 0 is real and does not contradict Metric‑T [18].
Φ⁰ = 2·min(%pos, %neg) = 2·min(1.0, 0.0) = 0 because all 672 point estimates share
one sign. That is a genuine property of this fixed feature (each electrode pair), not
an artifact of taking sign(|T|): it means on the AB-vs-CB contrast at the group-mean
level, no specification flips direction. It does not contradict [18] because the two
diagnostics average over different axes:
1. Φ fixes the feature and varies the analyst path (7 estimators × 16 preprocessing
× 6 inference). Metric‑T fixes the analytic path and varies the feature/connection.
A single feature can be direction-consistent across specs (Φ⁰ = 0) while different
features or a metric pair reverse — i.e. DC(M₁) − DC(M₂) can be non-zero even
when every spec on this one feature is same-signed. The axes of variation are
orthogonal, so agreement on one is silent about the other.
2. Differentiability: Φ⁰ answers "do the point-estimate directions agree within the
spec space?"; Metric‑T answers "does a metric's direction-consistency differ from
another metric's
across features?" A common positive bias on the AB cohort would
push every estimator's mean up together (all g > 0, Φ⁰ = 0) while leaving metric-to-
metric feature reversals (T) intact — exactly the "nuisance that biases every
specification in the same direction remains invisible" caveat the manuscript itself
states (§3.5).

So Φ⁰ = 0 is "all point estimates same-signed" directional agreement, which is a
statement about within-spec-space consensus on one feature, not a statement about
across-feature/across-metric reversal that [18] quantifies; it neither corroborates nor
refutes Metric‑T's T (which the manuscript, as noted in finding ADV‑004, never actually
reports).

Findings (grouped P2 → P3)

Detailed entries, evidence (file:line/path) and fixes are in artifacts/risk_register.json.

Legend for symbols used throughout: ✅ = identical at reported precision; ⚠ = discrepant/provenance; ➖ = not independently checkable (with reason); ≈ = close/explicable.

What this package already does right

What this audit does not guarantee

Artifacts on disk (under exec_check/)

Console log

==== START (status: OK) engine=Qwen/DeepSeek-V4-Flash rules=REPRO_STANDARDS.md date=2026-09-16 host=win32 py=3.8.10 numpy=1.24.4 scipy=1.10.1 pandas=2.0.3 matplotlib=3.7.5 package=v1.0 mode=live ====
[audit_reproducibility.py] 58 checks; 0 failed
[sobol3 self-test] S sum=1.000000000000; VR direct matches; I-constant S_I~1e-34 (pass; utf-8 console)
[fragile_geometry] R_robust=(-0.1949, +0.1172) numeric max-abs-err=3.44e-05; kappa=2.5 -> Phi_max=1.0 balance=True
[bia_confirmatory.py] all 7 index classifications + r* match Table 3; EXIT=0
[independent_recompute.R] R 4.6.1: 7/7 classifications OK; all verdicts & r* exact (diff=0.0000); EXIT=0
[bia_sobol3.py] Table 2 medians match shipped exactly; EXIT=0
[bia_sobol.py] Table 1 medians/percentiles match shipped; EXIT=0
[okimo_table6.py] Table 4 values match; schema differs (p_VR absent); EXIT=0
[eeg_confirmatory] common arm: 6 indeterminate + ciPLV equiv (matches); crossmix arm: 6 material, R_robust empty (matches)
[make_figures.py] fig1-4 pdf/eps/png written; EXIT=0
[delay_transition.py] analysis FAILED (delay_transition_raw.csv not shipped) -- expected per ADV-002
==== END (status: OK) ====

Generated 2026-09-16 by ReproAI audit (REPRO_STANDARDS.md). Source: REPRO_REPORT.md