A computational pipeline for structural and functional interpretation of disease-associated coding variants.
COMPASS is a computational framework designed to integrate whole-genome sequencing (WGS) data with experimentally determined and AI-predicted protein structures to systematically link disease-associated coding variants with functional structural regions and therapeutic targets. By combining sequence-based variant selection and structure-based patch scanning, COMPASS identifies disease-relevant substructures (functional hotspots) that accelerate genetics-driven drug discovery.
COMPASS contains three modules:
- COMPASS-Seq maps coding variants to transcript-specific amino-acid positions, resolves cases in which multiple variants affect the same residue, and performs gene-level association analysis on the selected variants.
- COMPASS-Struct scans a protein structure using residue-centered three-dimensional patches and tests the variants within each patch for association with the phenotype.
- COMPASS-Anno integrates functional annotations and literature evidence to assist interpretation of disease-associated structural hotspots.
COMPASS/
├── README.md
├── docs/
│ └── workflow.jpg
└── src/
├── coding_variants2amino_acids.R # COMPASS-Seq
├── patch_scanning.py # COMPASS-Struct driver
└── patch_scanning_p.R # Association tests used by COMPASS-Struct
A Docker image for COMPASS, which includes both R and Python environments with all COMPASS-related dependencies pre-installed, is available on Docker Hub. The image can be pulled using:
docker pull fengyn923/compass:0.1.0
COMPASS-Seq and the R component of COMPASS-Struct require R and the following packages:
gdsfmtSeqArraySeqVarToolsSTAARSTAARpipelineSTAARpipelineSummaryTxDb.Hsapiens.UCSC.hg38.knownGenedata.tablestringr
COMPASS-Struct additionally requires Python 3 and:
biopythonnumpypandasscipy
The Python driver runs patch_scanning_p.R through Rscript. Therefore, Rscript must be available on PATH, and patch_scanning_p.R must be accessible from the working directory where patch_scanning.py is launched.
Run commands from the directory containing the scripts, or provide paths that preserve access to patch_scanning_p.R.
cd src
Rscript coding_variants2amino_acids.R \
--category=missense \
--obj_nullmodel.file=/path/to/null_model.RData \
--chr=6 \
--gds.path=/path/to/chr6.annotated.gds \
--gene_name=CRIP3 \
--transcript=ENST00000274990 \
--seq.path=/path/to/sequences.csv \
--protein_sequence_selection_strategy=one-stepFor one-step, the selected variant table is written as:
CRIP3_ENST00000274990_onestep_variant.txt
Because patch_scanning.py currently calls patch_scanning_p.R by relative path, the safest invocation is from the src directory:
cd src
python patch_scanning.py \
--i_pdb_file=/path/to/CRIP3_structure.pdb \
--i_variant_file=/path/to/CRIP3_ENST00000274990_onestep_variant.txt \
--radius=20 \
--transcript=ENST00000274990 \
--gene_name=CRIP3 \
--obj_nullmodel=/path/to/null_model.RData \
--chr=6 \
--category=missense \
--gds_path=/path/to/chr6.annotated.gds \
--num_patch=10The final ranked hotspot table is written to:
CRIP3_ENST00000274990_patch.csv
| Parameter | Default | Description |
|---|---|---|
--obj_nullmodel.file |
— | Path to the RData file containing the pre-computed null model for the target phenotype. |
--category |
disruptive_missense |
Type of coding variants to analyze. Options: missense, disruptive_missense, ptv, plof. |
--chr |
1 |
Chromosome containing the target gene. |
--gds.path |
— | Annotated chromosome-level GDS file. |
--gene_name |
— | Gene symbol used to select variants. |
--transcript |
standard |
Ensembl transcript ID corresponding to the analyzed gene. The default setting is standard, which automatically selects the canonical transcript for analysis. |
--protein_sequence_selection_strategy |
one-step |
Strategy for resolving multiple variants at the same amino-acid position: one-step, stepwise, or best-subset-selection. |
--seq.path |
/data/sequences.csv |
Transcript/protein sequence catalogue CSV. |
--rare_maf_cutoff |
0.01 |
Rare-variant minor-allele-frequency cutoff passed to STAAR. |
--QC_label |
annotation/info/QC_label3 |
GDS node containing the variant QC label. |
--variant_type |
variant |
Variant type passed to STAARpipeline. |
--geno_missing_imputation |
mean |
Missing-genotype imputation method passed to STAARpipeline. |
--Annotation_dir |
annotation/info/FunctionalAnnotation |
GDS directory containing functional annotations. |
--Annotation_name_catalog.file |
/data/Annotation_name_catalog_V2.csv |
Annotation catalogue CSV. |
--Set_Based_Analysis.file |
/data/Set_Based_Analysis_Variant_List_20250224.R |
R script sourced at runtime; it must define Set_Based_Analysis_Variant_List. |
--Use_annotation_weights |
FALSE |
Whether STAAR annotation weights are used. |
--outfile |
— | Output file path. |
| Strategy | Behavior |
|---|---|
one-step |
For each amino-acid position mapped by multiple variant records, evaluates the association result after excluding each candidate and selects one representative variant for that position. |
stepwise |
Iteratively selects one representative variant for amino-acid positions mapped by multiple variant records, resolving one position at each iteration. |
best-subset-selection |
Evaluates every combination containing one candidate variant per amino-acid position and selects the combination with the smallest association P-value. Runtime increases rapidly with the number of positions and candidate variants. |
| Parameter | Default | Description |
|---|---|---|
--i_pdb_file |
— | Input PDB file containing the 3D structure of the target protein. |
--i_variant_file |
— | Variant file listing disease-associated variants (output from the sequence-based selection step). |
--radius |
15 |
Radius (in Å) used to define a 3D patch centered on each residue’s Cα atom. |
--transcript |
— | Ensembl transcript ID corresponding to the analyzed gene. |
--gene_name |
— | Name of the target gene for analysis. |
--obj_nullmodel |
— | Path to the RData file containing the pre-computed null model for the target phenotype. |
--chr |
— | Chromosome number containing the target gene. |
--category |
— | Variant category used in the patch-based association analysis. |
--gds_path |
— | Path to the annotated GDS file containing WGS variants for the specified chromosome. |
--num_patch |
10 |
Specifies the number of top-ranked structural hotspots (patches) with the lowest P-values to be reported. A value of 999 outputs all identified hotspots. |
--outfile |
— | Output file path. |
The output file is:
<gene>_<transcript>_patch.csv
Depending on the null model, the table contains:
| Column | Description |
|---|---|
central_id |
PDB residue number at the center of the patch. |
neighbors |
Comma-separated PDB residue numbers within the selected radius. |
variant_indices |
One-based row indices of variants from the input variant table. |
Chr |
Chromosome reported by the association analysis. |
Category |
Variant category. |
#SNV |
Number of SNVs included in the patch test. |
cMAC |
Cumulative minor allele count. |
STAAR-O or STAAR-B |
Association P-value used to rank patches. |
AA_Change |
Comma-separated transcript-specific amino-acid changes in the patch. |
COMPASS-Anno is used to provide functional interpretation for disease-associated structural hotspots identified by COMPASS. The interpretation is achieved by integrating functional context and literature-based evidence, with assistance from large language models, including ChatGPT and DeepSeek, and is accessible through our website (http://compassbio.net).
The current version is 0.1.0 (Jul 19, 2026).
Yannuo Feng*, Yihao Peng*, Shijie Fan*, Shijia Bian, Jingjing Gong, Qikai Jiang, Chang Lu#, Xihao Li#, Zilin Li#. Prioritizing druggable targets by mapping human disease-associated coding variants onto protein structures with COMPASS.
Contact: luc816@nenu.edu.cn, xihaoli@unc.edu, lizl@nenu.edu.cn
COMPASS is distributed under the GNU Affero General Public License v3.0. See LICENSE for details.
