Skip to content
li-lab-geneticsPublic

About

No description, website, or topics provided.

Resources

Stars

3 stars

Watchers

0 watching

Forks

Latest commit

 

History

31 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

COMPASS: COding variant Mapping onto Protein-structure-Anchored hotSpotS

A computational pipeline for structural and functional interpretation of disease-associated coding variants.

Description

COMPASS is a computational framework designed to integrate whole-genome sequencing (WGS) data with experimentally determined and AI-predicted protein structures to systematically link disease-associated coding variants with functional structural regions and therapeutic targets. By combining sequence-based variant selection and structure-based patch scanning, COMPASS identifies disease-relevant substructures (functional hotspots) that accelerate genetics-driven drug discovery.

Workflow Overview

COMPASS contains three modules:

  1. COMPASS-Seq maps coding variants to transcript-specific amino-acid positions, resolves cases in which multiple variants affect the same residue, and performs gene-level association analysis on the selected variants.
  2. COMPASS-Struct scans a protein structure using residue-centered three-dimensional patches and tests the variants within each patch for association with the phenotype.
  3. COMPASS-Anno integrates functional annotations and literature evidence to assist interpretation of disease-associated structural hotspots.

COMPASS_workflow

Repository structure

COMPASS/
├── README.md
├── docs/
│   └── workflow.jpg
└── src/
    ├── coding_variants2amino_acids.R  # COMPASS-Seq
    ├── patch_scanning.py              # COMPASS-Struct driver
    └── patch_scanning_p.R             # Association tests used by COMPASS-Struct

Requirements

Docker Image

A Docker image for COMPASS, which includes both R and Python environments with all COMPASS-related dependencies pre-installed, is available on Docker Hub. The image can be pulled using:

docker pull fengyn923/compass:0.1.0

Manual environment

COMPASS-Seq and the R component of COMPASS-Struct require R and the following packages:

  • gdsfmt
  • SeqArray
  • SeqVarTools
  • STAAR
  • STAARpipeline
  • STAARpipelineSummary
  • TxDb.Hsapiens.UCSC.hg38.knownGene
  • data.table
  • stringr

COMPASS-Struct additionally requires Python 3 and:

  • biopython
  • numpy
  • pandas
  • scipy

The Python driver runs patch_scanning_p.R through Rscript. Therefore, Rscript must be available on PATH, and patch_scanning_p.R must be accessible from the working directory where patch_scanning.py is launched.

Quick start

Run commands from the directory containing the scripts, or provide paths that preserve access to patch_scanning_p.R.

1. Run COMPASS-Seq

cd src

Rscript coding_variants2amino_acids.R \
  --category=missense \
  --obj_nullmodel.file=/path/to/null_model.RData \
  --chr=6 \
  --gds.path=/path/to/chr6.annotated.gds \
  --gene_name=CRIP3 \
  --transcript=ENST00000274990 \
  --seq.path=/path/to/sequences.csv \
  --protein_sequence_selection_strategy=one-step

For one-step, the selected variant table is written as:

CRIP3_ENST00000274990_onestep_variant.txt

2. Run COMPASS-Struct

Because patch_scanning.py currently calls patch_scanning_p.R by relative path, the safest invocation is from the src directory:

cd src

python patch_scanning.py \
  --i_pdb_file=/path/to/CRIP3_structure.pdb \
  --i_variant_file=/path/to/CRIP3_ENST00000274990_onestep_variant.txt \
  --radius=20 \
  --transcript=ENST00000274990 \
  --gene_name=CRIP3 \
  --obj_nullmodel=/path/to/null_model.RData \
  --chr=6 \
  --category=missense \
  --gds_path=/path/to/chr6.annotated.gds \
  --num_patch=10

The final ranked hotspot table is written to:

CRIP3_ENST00000274990_patch.csv

COMPASS-Seq

Parameters

Parameter Default Description
--obj_nullmodel.file — Path to the RData file containing the pre-computed null model for the target phenotype.
--category disruptive_missense Type of coding variants to analyze. Options: missense, disruptive_missense, ptv, plof.
--chr 1 Chromosome containing the target gene.
--gds.path — Annotated chromosome-level GDS file.
--gene_name — Gene symbol used to select variants.
--transcript standard Ensembl transcript ID corresponding to the analyzed gene. The default setting is standard, which automatically selects the canonical transcript for analysis.
--protein_sequence_selection_strategy one-step Strategy for resolving multiple variants at the same amino-acid position: one-step, stepwise, or best-subset-selection.
--seq.path /data/sequences.csv Transcript/protein sequence catalogue CSV.
--rare_maf_cutoff 0.01 Rare-variant minor-allele-frequency cutoff passed to STAAR.
--QC_label annotation/info/QC_label3 GDS node containing the variant QC label.
--variant_type variant Variant type passed to STAARpipeline.
--geno_missing_imputation mean Missing-genotype imputation method passed to STAARpipeline.
--Annotation_dir annotation/info/FunctionalAnnotation GDS directory containing functional annotations.
--Annotation_name_catalog.file /data/Annotation_name_catalog_V2.csv Annotation catalogue CSV.
--Set_Based_Analysis.file /data/Set_Based_Analysis_Variant_List_20250224.R R script sourced at runtime; it must define Set_Based_Analysis_Variant_List.
--Use_annotation_weights FALSE Whether STAAR annotation weights are used.
--outfile — Output file path.

Sequence-selection strategies

Strategy Behavior
one-step For each amino-acid position mapped by multiple variant records, evaluates the association result after excluding each candidate and selects one representative variant for that position.
stepwise Iteratively selects one representative variant for amino-acid positions mapped by multiple variant records, resolving one position at each iteration.
best-subset-selection Evaluates every combination containing one candidate variant per amino-acid position and selects the combination with the smallest association P-value. Runtime increases rapidly with the number of positions and candidate variants.

COMPASS-Struct

Main parameters

Parameter Default Description
--i_pdb_file — Input PDB file containing the 3D structure of the target protein.
--i_variant_file — Variant file listing disease-associated variants (output from the sequence-based selection step).
--radius 15 Radius (in Å) used to define a 3D patch centered on each residue’s Cα atom.
--transcript — Ensembl transcript ID corresponding to the analyzed gene.
--gene_name — Name of the target gene for analysis.
--obj_nullmodel — Path to the RData file containing the pre-computed null model for the target phenotype.
--chr — Chromosome number containing the target gene.
--category — Variant category used in the patch-based association analysis.
--gds_path — Path to the annotated GDS file containing WGS variants for the specified chromosome.
--num_patch 10 Specifies the number of top-ranked structural hotspots (patches) with the lowest P-values to be reported. A value of 999 outputs all identified hotspots.
--outfile — Output file path.

COMPASS-Struct output

The output file is:

<gene>_<transcript>_patch.csv

Depending on the null model, the table contains:

Column Description
central_id PDB residue number at the center of the patch.
neighbors Comma-separated PDB residue numbers within the selected radius.
variant_indices One-based row indices of variants from the input variant table.
Chr Chromosome reported by the association analysis.
Category Variant category.
#SNV Number of SNVs included in the patch test.
cMAC Cumulative minor allele count.
STAAR-O or STAAR-B Association P-value used to rank patches.
AA_Change Comma-separated transcript-specific amino-acid changes in the patch.

COMPASS-Anno

COMPASS-Anno is used to provide functional interpretation for disease-associated structural hotspots identified by COMPASS. The interpretation is achieved by integrating functional context and literature-based evidence, with assistance from large language models, including ChatGPT and DeepSeek, and is accessible through our website (http://compassbio.net).

Version

The current version is 0.1.0 (Jul 19, 2026).

Authors

Yannuo Feng*, Yihao Peng*, Shijie Fan*, Shijia Bian, Jingjing Gong, Qikai Jiang, Chang Lu#, Xihao Li#, Zilin Li#. Prioritizing druggable targets by mapping human disease-associated coding variants onto protein structures with COMPASS.

Contact: luc816@nenu.edu.cn, xihaoli@unc.edu, lizl@nenu.edu.cn

License

COMPASS is distributed under the GNU Affero General Public License v3.0. See LICENSE for details.

About

No description, website, or topics provided.

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages