Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CITATION.cff
Original file line number Diff line number Diff line change
Expand Up @@ -8,5 +8,5 @@ authors:
repository-code: "https://github.com/oncologylab/survscope"
url: "https://oncologylab.github.io/survscope/"
license: MIT
version: 0.4.1
version: 0.4.2
date-released: 2026-09-21
8 changes: 4 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,9 +29,9 @@ TCGA provides overall survival (OS), disease-specific survival (DSS), progressio

## Read the results carefully

A curve estimates the fraction of patients who remain event-free over time. The legend gives patients (`n`) and events (`e`) in each group. The p-value compares the curves; the hazard ratio compares higher with lower expression. These are unadjusted associations, not proof that a gene causes a difference or predicts an individual's outcome.
A curve estimates the fraction of patients who remain event-free over time. The legend gives patients (`n`) and observed events (`e`) in each group. Patients without an observed event are censored, so `e ≤ n`. Double-click a whole legend entry to edit its displayed text; the calculated counts remain in the analysis results. The p-value compares the curves; the hazard ratio compares higher with lower expression. These are unadjusted associations, not proof that a gene causes a difference or predicts an individual's outcome.

Comparing selected lower and upper percentile groups leaves out the middle patients and can increase uncertainty. Trying several genes or group definitions adds multiple comparisons; the displayed q-value adjusts only for the available outcomes **within one analysis**. Choose comparisons for a scientific reason and report what you explored. [Learn to read the plot](docs/user-guide.md#understand-the-figure).
Comparing selected lower and upper percentile groups leaves out the middle patients and can increase uncertainty. Trying several genes or group definitions adds multiple comparisons; the displayed q-value adjusts only for the available outcomes **within one analysis**. With one tested outcome, as in current CPTAC analyses, `q = p`; the figure shows p alone. Choose comparisons for a scientific reason and report what you explored. [Learn to read the plot](docs/user-guide.md#understand-the-figure).

JavaScript calculations are checked against Python and independently executed R `survival`. Group membership, event counts, and risk counts agree exactly in the validation suite. Numerical tolerances and the preserved PAAD reference estimates are documented in [methods and validation](docs/methods.md#validation).

Expand All @@ -43,11 +43,11 @@ TCGA analyses cite **TCGA-CDR** for survival outcomes, **UCSC Xena** for data di

## Use Python or the command line

Install the tested [GitHub release](https://github.com/oncologylab/survscope/releases/tag/v0.4.1):
Install the tested [GitHub release](https://github.com/oncologylab/survscope/releases/tag/v0.4.2):

```bash
python -m pip install \
https://github.com/oncologylab/survscope/releases/download/v0.4.1/survscope-0.4.1-py3-none-any.whl
https://github.com/oncologylab/survscope/releases/download/v0.4.2/survscope-0.4.2-py3-none-any.whl
survscope plot --gene SRD5A1 --cohort PAAD --format pdf svg png --outdir plots
survscope plot --gene TP53 --cohort CPTAC-3-LUAD \
--grouping percentile_groups --lower-percent 25 --upper-percent 25 --json --outdir plots
Expand Down
2 changes: 1 addition & 1 deletion docs/cptac.md
Original file line number Diff line number Diff line change
Expand Up @@ -111,7 +111,7 @@ CPTAC OS receives a **caution** badge because follow-up completeness varies.
The TCGA-CDR quality recommendations are not assigned to CPTAC. DSS, PFI and DFI
are null arrays marked **unavailable**, with explicit unavailable figure panels.
No recurrence field is relabeled as a TCGA-CDR endpoint. BH adjustment includes
only finite endpoint p-values, so CPTAC's OS q-value equals its OS p-value.
only finite endpoint p-values, so CPTAC's OS q-value equals its OS p-value. The figure shows p alone for this single-test case; both values remain in analysis JSON.

Solid-tumor cohorts use GDC `Primary Tumor` sample metadata. AML accepts primary
bone marrow or peripheral-blood cancer specimens, preferring bone marrow when
Expand Down
2 changes: 1 addition & 1 deletion docs/development.md
Original file line number Diff line number Diff line change
Expand Up @@ -82,7 +82,7 @@ This reuses checksummed compact TCGA assets and builds CPTAC. Leave `tcga_data_v

## Deployment and software publishing

Software (`v0.4.1`, …) and immutable data (`data-vYYYY.MM.DD`) have independent versions. No data rebuild is required for editor or analysis-software updates.
Software (`v0.4.2`, …) and immutable data (`data-vYYYY.MM.DD`) have independent versions. No data rebuild is required for editor or analysis-software updates.

Pages deployment starts after a successful main-branch CI run, a successful main-branch data-release workflow, or an explicit main-branch dispatch. It checks out that triggering commit, runs Python and browser checks, verifies the full data release, and validates statistics across its deployed cohort catalog. Deployment requires passing tests and a complete built site below **891,289,600 bytes (850 MiB)**. The release archive stays immutable; fonts/editor code count toward the final site budget.

Expand Down
Binary file modified docs/images/citations.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file modified docs/images/comparison.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file modified docs/images/editor.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
4 changes: 3 additions & 1 deletion docs/methods.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,7 +36,7 @@ The unadjusted binary Cox proportional-hazards model compares higher with lower

To preserve the published PAAD values, the Python Cox solver retains SciPy's bounded minimization, initially on log-HR [−8, 8], with `xatol=1e-5` and a maximum of 500 iterations. JavaScript uses the matching solver and NumPy reduction order. Before fitting, both implementations check mixed-group event risk sets for identifiability and separation. A finite optimum outside the initial interval triggers bracket expansion, up to ±64. Empty groups, no events, no information, separation, or failed convergence return `NA` with a reason; a boundary value is not presented as an estimated HR. Complete separation can occur even when both groups have events.

Benjamini–Hochberg q-values adjust only finite log-rank p-values across the four outcomes for the current gene, cohort, and grouping. Hiding a panel does not change that family. CPTAC has one available endpoint, so its OS q-value equals its p-value. This adjustment does **not** account for exploration across genes, cohorts, or grouping choices.
Benjamini–Hochberg q-values adjust only finite log-rank p-values across the four outcomes for the current gene, cohort, and grouping. Hiding a panel does not change that family. CPTAC has one available endpoint, so its OS q-value equals its p-value. The figure omits this redundant q label when only one finite endpoint p-value exists; analysis JSON retains the calculated q. Several-outcome BH adjustment can also leave some p-values unchanged, and display rounding can conceal small differences. This adjustment does **not** account for exploration across genes, cohorts, or grouping choices.

## Confidence bands, censor marks, and risk tables

Expand Down Expand Up @@ -65,3 +65,5 @@ The acceptance limits are:
Undefined estimates must agree in availability. The R Cox tolerance is deliberately distinct: preserving the original bounded-solver reference means it will not match every final decimal of R's more tightly converged estimate. In the full release check, all counts and memberships matched exactly; maximum log-HR differences were 2.59e-9 (JavaScript) and 2.14e-6 (R), and maximum Cox p-value differences were 4.17e-9 and 7.89e-6, respectively. The [validation report](validation-2026.09.21.json) records observed errors and coverage.

The original SRD5A1/PAAD tests retain its endpoint sample/event counts, exact median membership, p/q/HR values, blue/red default, and 6.8-inch figure. Browser checks also exercise direct editing, undo, projects, optional overlays, physical SVG/PDF dimensions, PNG pixels, embedded fonts, small screens, and same-origin runtime requests. [Development instructions](development.md) explain how to reproduce these checks.

Legend entries are editable presentation text. Editing a displayed name or count does not change group membership, curves, statistics, or the counts in analysis JSON and citation methods. Each outcome has its own override; resetting it or creating a new analysis restores automatic data labels.
4 changes: 2 additions & 2 deletions docs/publishing.md
Original file line number Diff line number Diff line change
@@ -1,12 +1,12 @@
# Python package publishing

The [v0.4.1 GitHub release](https://github.com/oncologylab/survscope/releases/tag/v0.4.1)
The [v0.4.2 GitHub release](https://github.com/oncologylab/survscope/releases/tag/v0.4.2)
contains the tested wheel and source distribution. The website and immutable
data release are already public. Install the GitHub wheel directly:

```bash
python -m pip install \
https://github.com/oncologylab/survscope/releases/download/v0.4.1/survscope-0.4.1-py3-none-any.whl
https://github.com/oncologylab/survscope/releases/download/v0.4.2/survscope-0.4.2-py3-none-any.whl
```

The first PyPI upload was rejected with `invalid-publisher`: the signed GitHub
Expand Down
14 changes: 9 additions & 5 deletions docs/user-guide.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,9 +45,9 @@ A Kaplan–Meier curve starts at 1 (100%). It steps down when an event occurs. A

CPTAC currently supports OS only. Its remaining panels say **Endpoint unavailable**. TCGA quality labels reproduce the source's recommendations; a **caution** or **not recommended** outcome needs particular care. CPTAC has a separate caution about follow-up completeness. Very small groups may have no usable group comparison.

- **n** is the number of patients in a group; **e** is the number of events.
- **n** is the number of patients with usable follow-up in a group; **e** counts patients with an observed event. For OS, that event is death. The other patients are censored: follow-up ended without the event being observed. Therefore **e ≤ n**; equality is possible when every included patient has an event.
- **p** is the two-sided log-rank p-value. It concerns the difference between these two curves, not whether the finding will reproduce elsewhere.
- **q** adjusts the available endpoint p-values for this one gene, cohort, and comparison. It does not correct for trying many genes, cohorts, or cutoffs.
- **q** adjusts the available endpoint p-values for this one gene, cohort, and comparison. With only one tested outcome (current CPTAC analyses), **q = p**, so the figure shows p alone. With several outcomes, some BH-adjusted values can also equal p, and rounded values may look equal. Hiding panels does not change the adjustment. It does not correct for trying many genes, cohorts, or cutoffs; analysis JSON keeps the full-precision values.
- **HR** is the unadjusted hazard ratio for higher versus lower expression. Above 1 indicates a higher modeled event hazard in the higher-expression group. Interpretation assumes proportional hazards; the application does not test this assumption or adjust for stage, treatment, age, or other factors.
- **NA** means the statistic cannot be estimated. Empty groups, no events, lack of overlapping risk sets, or complete separation can cause this. The available curves may still be shown.

Expand All @@ -61,9 +61,9 @@ These three additions are off initially, preserving the original appearance. Res

## Edit the figure

Your figure is ready to edit as soon as it appears. Double-click a title, axis label, group name, or note to type directly on the figure. Select individual characters to apply bold, italic, superscript, or subscript using the floating text toolbar. Enter adds a line; Escape or clicking outside finishes the edit. Saving or exporting also finishes an active text edit.
Your figure is ready to edit as soon as it appears. Double-click a title, axis label, complete legend entry, or note to type directly on the figure. Select individual characters to apply bold, italic, superscript, or subscript using the floating text toolbar. Enter adds a line; Escape or clicking outside finishes the edit. Saving or exporting also finishes an active text edit.

The editing commands use familiar icons: a pointer for Selection, a T for Type, a hand for panning, a disk for saving, and curved arrows for Undo/Redo. Hover over an icon, or focus it with Tab, to see its name and keyboard shortcut. Alignment, distribution, copying, annotation order, and text formatting use the same icon controls. Unavailable actions are dimmed. The information icon beside **Selected item** explains direct editing.
The editing commands use familiar icons: a pointer for Selection, a T for Type, a hand for panning, a disk for saving, and curved arrows for Undo/Redo. Hover over an icon, or focus it with Tab, to see its name and keyboard shortcut. Alignment, distribution, copying, annotation order, and text formatting use the same icon controls. Unavailable actions are dimmed. The compact toolbar at the top of Properties has alignment and annotation commands; its number badge shows how many objects are selected.

The workspace fills wide displays. **Analysis** sits on the left and **Properties / Layers** on the right; use the rail buttons to open or close them. Drag the panel boundaries to adjust their widths. On smaller screens, the panels open as drawers so the figure keeps its space. **Reset workspace** restores the panel sizes and Fit view without changing your figure.

Expand All @@ -90,7 +90,7 @@ The editor offers:
- **Size and appearance:** width and height in inches, font family, text scale, colors, line width, and solid/dashed/dotted curves. Titles and labels also have individual text and style controls; Enter in a numeric field applies the value.
- **Outcomes and layout:** show selected panels, arrange a grid or row/column, and change their order. This changes presentation only; q-values still reflect all valid outcomes from the analysis.
- **Axes:** show months, years, or days; set a displayed time range and tick spacing; zoom into a survival range. Axis limits do not truncate follow-up or refit the model. Extremely dense requested ticks are thinned to keep the figure responsive.
- **Legend and annotations:** rename the lower/higher group labels and add notes, lines, or arrows. Text supports multiple lines. Computed statistical values update with the analysis and cannot be overwritten in the editor.
- **Legend and annotations:** double-click any complete legend entry, including its `n=…` and `e=…`, to edit and format all its text. Each outcome has separate entries. Legend edits change the figure only; the analysis JSON, curves, and citations retain the original calculated counts. **Reset selected item** restores the automatic entry. The lower/higher group-name fields apply names across all outcomes and restore their automatic counts. Add notes, lines, or arrows, and use multiple lines when needed. Calculated p, q, and HR values remain linked to the analysis.

Use **Fit**, **100%**, the zoom buttons, and the **Hand tool** to inspect details without changing export dimensions. With focus on the figure, arrow keys move the selected item by one point; hold Shift for ten points. **Undo/Redo** also respond to Ctrl/Cmd+Z and Ctrl/Cmd+Shift+Z. **Reset figure** returns to the original appearance and can itself be undone.

Expand Down Expand Up @@ -141,6 +141,10 @@ Save the project to keep this citation snapshot with the figure. [Full reference

## Common questions

**Why is n larger than e?** A patient can contribute follow-up without having an observed event. For example, n=100 and e=30 means 30 events and 70 censored observations. n and e can be equal; e cannot exceed n.

**Why are p and q equal in CPTAC?** Current CPTAC data provide overall survival only. Benjamini–Hochberg adjustment of one test leaves its p-value unchanged. The figure omits the duplicate q label, while JSON retains both values. The information icon next to the q control explains how many outcomes were tested.

**Why do the group counts differ from the cohort count?** Missing expression, missing outcomes, or nonpositive follow-up time exclude a patient from that outcome. Custom percentile comparisons also leave out the middle expression values.

**Why did a percentile not split the group evenly?** Equal expression values are kept together. Percentiles define expression thresholds, not an arbitrary ordering of people with the same value.
Expand Down
2 changes: 1 addition & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"

[project]
name = "survscope"
version = "0.4.1"
version = "0.4.2"
description = "Reproducible TCGA and CPTAC Kaplan-Meier survival plots from compact static data"
readme = "README.md"
license = "MIT"
Expand Down
2 changes: 1 addition & 1 deletion src/survscope/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@
"search_genes",
]

__version__ = "0.4.1"
__version__ = "0.4.2"


def available_cohorts(store: DataStore | None = None) -> list[dict]:
Expand Down
11 changes: 6 additions & 5 deletions src/survscope/plotting.py
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,7 @@ def create_figure(analysis: SurvivalAnalysis):
"""Create but do not save the canonical 6.8-inch figure."""
apply_plot_style()
figure, axes = plt.subplots(2, 2, figsize=(6.8, 6.8))
show_q = sum(np.isfinite(result.logrank_p) for result in analysis.endpoints.values()) > 1
for axis, endpoint in zip(axes.flat, ENDPOINTS, strict=True):
result = analysis.endpoints[endpoint]
if result.quality == "unavailable":
Expand All @@ -68,11 +69,11 @@ def create_figure(analysis: SurvivalAnalysis):
linewidth=2.1,
label=f"{label} n={curve.n}, e={curve.events}",
)
annotation = (
f"p={format_p(result.logrank_p)} q={format_p(result.logrank_q)}\nHR={result.cox_hr:.2f}"
if np.isfinite(result.cox_hr)
else f"p={format_p(result.logrank_p)} q={format_p(result.logrank_q)}\nHR=NA"
)
annotation = f"p={format_p(result.logrank_p)}"
if show_q:
annotation += f" q={format_p(result.logrank_q)}"
hr = f"{result.cox_hr:.2f}" if np.isfinite(result.cox_hr) else "NA"
annotation += f"\nHR={hr}"
axis.text(
0.97,
0.06,
Expand Down
7 changes: 5 additions & 2 deletions tests/test_cptac.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@
import numpy as np
import pytest

from survscope import analyze
from survscope import analyze, plot
from survscope.builder import SourceText, build_release
from survscope.cptac import (
CPTAC_COHORTS,
Expand Down Expand Up @@ -331,7 +331,7 @@ def get(path, **params):
client.search("cases", {}, "case_id")


def test_live_cptac_fixture_analysis_and_reference_contract():
def test_live_cptac_fixture_analysis_and_reference_contract(tmp_path):
store = DataStore(base=CPTAC_FIXTURE, data_version="2026.09.18", cache=False)
result = analyze("SRD5A1", "CPTAC-3-PAAD", store=store)
assert result.filename_stem == "SRD5A1_CPTAC_3_PAAD_KM_survival"
Expand All @@ -342,6 +342,9 @@ def test_live_cptac_fixture_analysis_and_reference_contract():
assert result.endpoints["OS"].logrank_q == result.endpoints["OS"].logrank_p
assert all(result.endpoints[ep].quality == "unavailable" for ep in ("DSS", "PFI", "DFI"))
assert all(np.isnan(result.endpoints[ep].cox_hr) for ep in ("DSS", "PFI", "DFI"))
figure = plot(result, formats=("svg",), output_dir=tmp_path).paths[0].read_text()
assert "p=0.21" in figure
assert "q=" not in figure


def test_release_validation_detects_tampering_and_missing_coverage(tmp_path):
Expand Down
1 change: 1 addition & 0 deletions tests/test_reference.py
Original file line number Diff line number Diff line change
Expand Up @@ -82,6 +82,7 @@ def test_reference_figure_contract(store, tmp_path: Path):
assert 'width="489.6pt"' in svg
assert "SRD5A1 TCGA-PAAD survival" in svg
assert "PanCanAtlas TCGA-CDR" in svg
assert "p=0.0018 q=0.0069" in svg
assert outputs.paths[2].stat().st_size > 20_000


Expand Down
4 changes: 2 additions & 2 deletions web/package-lock.json

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion web/package.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"name": "survscope-web",
"private": true,
"version": "0.4.1",
"version": "0.4.2",
"type": "module",
"scripts": {
"dev": "vite",
Expand Down
Loading
Loading