Transfer learning for genomic prediction in underrepresented populations
The impact of meta-analysis and PRS-CSx
The above experiments restricted PRS model input variants to those discovered in the UKB European population, excluding any trait-associated variants unique to BBJ samples. To extend the analyses to capture these, we created two additional prediction methods.
First, we ran GWAS on each BBJ sample size. The first new method performed a cross-population meta-analysis using the full UKB GWAS and the sample-size-specific BBJ GWAS to identify candidate variants, and then fit an elastic net on those variants. The second method used PRS-CSx to combine the two sets of GWAS summary statistics.
By tracking the net performance gain of meta-analysis and PRS-CSx across varying discovery sample sizes, we observed differences in performance across BBJ sample sizes.
The influence of meta-analysis is low for conserved traits, largely due to reduced statistical power in the much smaller BBJ GWAS sample sizes. However, for population-specific traits like HDL and LDL, and to a lesser extent blood glucose, meta-analysis substantially outperforms single-population discovery. The gains are primarily due to modifying the elastic net variant input: including UKB European samples during training slightly improved prediction when using 10,000 or fewer BBJ samples. This improvement was not seen with larger BBJ samples.
Because PRS-CSx dynamically weights population-specific models, its performance is theoretically less sensitive to conserved vs population-specific traits. However, we observed that the model requires more data than elastic net models to perform well. For target sample sizes under 25k, PRS-CSx performs worse than the strongest corresponding elastic net model in all phenotypes except BMI. As sample sizes approached 100k, PRS-CSx matched or exceeded the best performing model across all phenotypes except blood glucose.