Using UK Biobank data across three phenotypically distinct exposures — circulating amino acids, BMI, and major depressive disorder (MDD) — a systematic evaluation of linkage disequilibrium (LD) clumping thresholds reveals that the near-universal default parameters (r² <0.001, 10,000 kb window) routinely applied in Mendelian randomization (MR) studies are suboptimal. For amino acids and BMI, more stringent LD thresholds (r²=0.00001) with larger genomic windows strengthened instruments, while MDD — a highly polygenic binary trait — responded better to smaller windows with equivalent r² stringency. Crucially, simply adding more SNPs did not reliably improve instrument quality, exposing a pleiotropy-strength trade-off that default pipelines ignore.
This finding matters because MR has become a cornerstone of causal inference in nutritional, metabolic, and psychiatric epidemiology, directly shaping clinical and public health recommendations. Flawed instrument selection introduces weak-instrument bias and inflates pleiotropy risk, potentially generating spurious causal conclusions that propagate through systematic reviews and policy. The authors' proposed empirical framework — evaluating both F-statistics and variance explained (R²) per exposure type — is a practical corrective that could meaningfully improve reproducibility across the field.
Limitations deserve emphasis: this is a methodological paper, not a novel health discovery, and generalizability beyond the three tested phenotypes is unproven. As a preprint posted on medRxiv and not yet peer-reviewed, the proposed framework requires independent validation before adoption as a community standard. Nonetheless, the work is a timely and potentially practice-changing contribution to causal epidemiology methodology.