Skip to content

About

A repository for integrating and evaluating genomic and proteomic risk scores for multi-ancestry populations.

Resources

Stars

0 stars

Watchers

0 watching

Forks

 
 

Repository files navigation

Genomic and Proteomic Risk Score Integration

A repository for integrating and evaluating genomic and proteomic risk scores for multi-ancestry populations.

Overview

An imputation and combined risk score framework for evaluating individual risk of developing dichotomous disease traits, as well as prediction performance for continuous traits. We benchmark and select from various imputation approaches using UK Biobank plasma proteomics and apply the selected approach to protein expression data across ancestries. Ensemble learning on five protein risk prediction models estimates the optimal weights for obtaining protein risk scores (ProRS). The framework incorporates data pre-processing and normalization steps to further enhance the prediction performance of the combined polygenic risk score (PRS) and ProRS model.

Evaluation of the PRS and ProRS models is completed in the following steps:

  • ProRS importance scores
  • PRS + ProRS prediction performance
  • PRS and ProRS mediation
  • PRS and ProRS prognostic and diagnostic performance

Figure 1

Package and software requirements

All analyses were run within R version 4.5.0. We recommend using the Intel Math Kernel Library or equivalent to improve computational speed. For PRS scoring, we used PLINK 2.00.

R packages:

  • General: dplyr (version 1.1.4), ggplot2 (version 4.0.0), ggpubr (version 0.6.1), stringr (version 1.5.1), tibble (version 3.3.0), tidyr (version 1.3.1)
  • Imputation: mice (version 3.18.0), missForest (version 1.5), pcaMethods (version 2.00), PEMM (version 1.0)
  • Risk score modeling and evaluation: caret (version 7.0-1), glmnet (version 4.1-10), mlr (version 2.19.3), pROC (version 1.18.5), randomForest (version 4.7-1.2), regmedint (version 1.0.1), RISCA (version 1.0.7), SuperLearner (version 2.0-29), xgboost (version 3.2.1.1)

To install a specific package version,

install.packages("remotes")
remotes::install_version("dplyr", version = "1.1.4")
#or
devtools::install_version("dplyr", version = "1.1.4")

Contents

  1. Data pre-processing: located in the Data_preprocessing folder, individual phenotypes, protein expression data, batch ID and visit from the UK Biobank are loaded and aggregated into one dataset in Protein_preprocess.R. For obtaining trait status from Field IDs, please refer to UKB_LargePhenotypesPull.R and for time of disease diagnosis, the script can be found in UKB_Phenotypes_RollingScript.R. Further details on data formatting can be found in the scripts.

  2. Imputation benchmarking: found in the Imputation_benchmarking folder, five imputation approaches are benchmarked on protein expression data from the UKB. After adjusting for batch effects, the proportion of missingness or missingness rate is estimated and NA assigned according to this rate for each protein in Imputation_data_save.R. The data with assigned missingness is loaded into downstream tasks within this same folder to benchmark the performance of imputation algorithms.

  3. PRS Scoring: in the PRS_scoring folder, effect size estimates for each trait from the PGS Catalog are applied to genotype data for calculating PRS of all individuals. In PRS_Scoring_Script.R, variant lists are created by matching SNP identifiers between UKB and PGS Catalog datasets and following scoring using PLINK, risk scores are combined across chromosomes. Shell scripts for subsetting variant ranges and completing variant scoring are found in bgen_ranges and corresponding trait folders, respectively.

  4. ProRS Scoring: the ProRS_scoring folder contains the shell script for running each step of protein risk score estimation. For a given trait, the output from protein imputation is saved for downstream tasks such as machine learning model fitting and prediction, and ensemble learning using SuperLearner. ProRS model components are saved for visualizing protein importance scores in figures.

  5. PRS+ProRS integration and mediation: for each trait, combining and evaluating the PRS and ProRS prediction estimates are completed in scripts named PRS_ProRS_[Trait].R (with pseudo R-squared equivalents for disease traits), or PRS_ProRS_[Trait]_strat.R in male/female stratified analyses. The prediction performance of each component used as input for the SuperLearner is estimated in ProRS_components_[Trait].R. Mediation analyses for quantifying the proportion of PRS effects mediated by ProRS are completed in scripts titled PRS_ProRS_[Trait]_mediation.R.

  6. Prognostic and diagnostic analyses: PRS and ProRS model performance across different time intervals between time of biomarker collection and disease diagnosis are evaluated in Prog_diag_ProRS.R, or Prog_diag_ProRS_strat.R for female/male stratified analyses.

The codes for creating the figures in the manuscript are located in the Figures folder.

About

A repository for integrating and evaluating genomic and proteomic risk scores for multi-ancestry populations.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages