Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
57 changes: 57 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,63 @@ All notable changes to Tessera are recorded here. The format follows

## [Unreleased]

### Fixed

- **The MAFFT backend ignored strand.** A genome, or one contig of a draft assembly, on the
opposite strand to the backbone was aligned as given and came out at chance-level identity
(about 0.40 against 0.97 for the same genome in forward orientation), which the scan then
read as a divergent region. MAFFT now runs with `--adjustdirection`.
- **`reassort` never applied its alignment-fraction filter.** The threshold was written as a
fraction (0.5) and compared with skani's percentage, so a tip aligning over a fifth of a
segment could outrank a full-length match and become the segment's nearest strain. The
threshold is now 50 %.
- **`fill-references` did not align the last round's downloads.** On a `--max-rounds` exit the
final downloads were counted in `fill_summary.tsv`, the report and `lineages.tsv` but were
absent from `panel.msa.fasta`, so detection ran without them. One more alignment
(`final.msa.fasta`) is now built from them, without a further search.
- **`fill-references --curate` did nothing unless a round both found a gap and downloaded a
reference.** A supplied collection that contains a whole-genome sibling of the query has
no coverage gap, so the loop converged before curation ran. Curation now runs before the
first alignment for a supplied collection and before each later build for that round's
downloads. A panel seeded from scratch is unchanged: seeding has its own sibling filter.
One limit remains: the sibling test is anchored on the query's closest whole-genome
relative, so when a sibling is itself the closest genome in the collection it becomes
that anchor and is kept (further siblings are removed). The log names the anchor chosen.
- **`fill-references --curate --reference X` could delete X** and fail the next round with
"Reference 'X' not found". The given reference is now never removed by curation (it is
listed as `sibling-kept` / `redundant-kept` where it would have been). The sibling test
stays anchored on the query's closest relative, not on the reference.
- **A run could delete an existing `<output>/collection` directory that was not its own.**
`detect`, `fill-references`, `build-panel` and `curate-panel` clear that directory to hold
their working copy of the references. If it already exists, is not empty, and no earlier
run left its files beside it, the run now stops with an error and nothing is deleted.
- **`find-references --download <collection> --curate` deleted genomes that were already in
the collection.** Curation now removes only what the run downloaded; genomes already
present are reported as `sibling-kept` / `redundant-kept` in `panel_lineages.tsv`. A
downloaded sibling can no longer be picked as the curation backbone either.
- **`reassort --scan-segments` could write outside the output directory.** A segment record
named `..` resolved the scan directory to the parent of the output and replaced its
`collection/`. Segment names are sanitised by one shared helper, and two names that
sanitise alike no longer share a scan directory.
- **MAF-based backends (`sibeliaz`, `cactus`) could drop rows and columns.** A genome placed in
no alignment block had no row at all; it is now an all-gap row. It is named in a warning,
as is a genome aligned only in blocks that do not include the backbone. With
`sibeliaz`, a multi-contig backbone was laid out in sorted-name order and lost any contig
that no block covered, shifting later coordinates; it is now laid out as its FASTA file is,
with uncovered contigs kept as gap columns.
- **`progressivemauve` named rows after the resolved input path.** Symlinked collections,
unrecognised extensions and whitespace in a path gave wrong or duplicate row names, and a
genome whose link target shared the backbone's name was dropped. Rows are now named from
the staged genome label. (Checked against a simulated XMFA; the tool itself was not
available.)
- **Whitespace in sequence lines was read as sequence.** A reference saved with trailing
spaces, tabs or Windows line endings produced a ragged alignment. Such a genome is now
aligned from a cleaned copy (a warning names it), and the FASTA reader ignores whitespace
in sequence lines.
- A header line `> ` (no name), a truncated MAF block, an XMFA that does not list the
reference and an empty MAFFT result are reported as input/output errors that name the
file, instead of "Unexpected error".

## [1.2.0] - 2026-10-01

### Fixed
Expand Down
17 changes: 14 additions & 3 deletions docs/aligners.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ with `--aligner` and tune with repeatable `--aligner-arg key=value`.
| Backend | Best for | Notes |
|---|---|---|
| `sibeliaz` (default) | Moderately divergent genomes, including rearrangements | Installs cleanly via conda; `kmer`, `abundance`, `bubble`, `filtermemory` |
| `mafft` | Similar, largely collinear genomes | True base-level alignment, the canonical input for the window method; adds a fragmented query with `--addfragments`. `maxiterate`, `retree`, `op`, `ep`, `sixmerpair` |
| `mafft` | Similar, largely collinear genomes | True base-level alignment, the canonical input for the window method; adds a fragmented query with `--addfragments` and reorients reverse-strand genomes or contigs (`--adjustdirection`). `maxiterate`, `retree`, `op`, `ep`, `sixmerpair` |
| `minimap2` | Speed and assembly/contig queries | Fast assembly-to-reference projection; `preset` (default `asm20`, e.g. `asm10` for closer genomes) |
| `progressivemauve` | Genomes with large rearrangements/inversions | Tolerant but slow, heavy, and not available as a conda build on all platforms; `seed_weight`, `single` |
| `cactus` | Same-species pangenomes | Resource heavy (Toil/containers) |
Expand All @@ -17,8 +17,19 @@ with `--aligner` and tune with repeatable `--aligner-arg key=value`.
data, reproduces `progressivemauve`'s recombination coordinates. For very similar,
collinear genomes `mafft` gives the most faithful base-level signal and `minimap2`
the fastest run (and the best fit for a fragmented query); `progressivemauve` remains
an option for genomes with large rearrangements. Reference-anchored backends drop
material inserted relative to the backbone; `mafft` keeps it as a true alignment.
an option for genomes with large rearrangements. Every backend reports the alignment in
backbone coordinates, so material inserted relative to the backbone is dropped -- `mafft`
included, which runs with `--keeplength`.

A multi-contig backbone is laid out as the concatenation of its contigs in the order of
its FASTA file (`sibeliaz`, `mafft`, `minimap2`, `progressivemauve`); a contig that nothing
aligns to keeps its columns, as gaps. The `cactus` backend lays contigs out in sorted-name
order. A genome the aligner cannot place against the backbone at all is kept as an all-gap
row and named in a warning, so the alignment always has one row per input genome.

Genomes are staged before alignment. A FASTA with whitespace inside its sequence lines
(trailing spaces, tabs, Windows line endings) is aligned from a cleaned copy and reported
in a warning; the input file itself is not modified.

Examples:

Expand Down
27 changes: 22 additions & 5 deletions docs/reference-panels.md
Original file line number Diff line number Diff line change
Expand Up @@ -105,7 +105,10 @@ It stops when the gaps close, when no new reference can be found, or when a roun
longer improves the worst gap (so a genuinely hypervariable region is reported, not
chased forever). Each round is recorded in `filled/fill_summary.tsv`, and the query's
own record is auto-excluded from its FASTA header. This needs an aligner and Entrez
Direct, and rebuilds the alignment every round.
Direct, and rebuilds the alignment every round. If the last permitted round
(`--max-rounds`) still downloaded references, one more alignment (`final.msa.fasta`) is
built from them without a further search, so `panel.msa.fasta` holds every reference the
run reports.

Omit `--collection` to **start fresh** with no suggested references: the first round
seeds the collection from an NCBI search, then the loop fills the remaining gaps as
Expand Down Expand Up @@ -277,9 +280,23 @@ SARS-CoV-2 sublineages (<1 % apart) alike. The curated `curated/collection/` and
`panel_lineages.tsv` (each reference's role and ANI/coverage) are written; rebuild with
`tessera msa` then `tessera recomb`.

The same curation runs inside the fill loop with `fill-references --curate`, which
keeps the growing panel diverse and sibling-free each round and adds a "Reference
panel" section to the report. Both need skani (and skDER for dereplication):
The same curation runs inside the fill loop with `fill-references --curate`, and adds a
"Reference panel" section to the report. It runs before an alignment is built whenever the
working collection holds genomes that have not been through it: the collection you
supplied (before round 1), and each round's downloads (before the next build). A panel
seeded from scratch is not curated before round 1, because seeding applies its own sibling
filter. With `--reference`, that genome is never removed by curation; the sibling test
stays anchored on the query's closest whole-genome relative, which the log names.

One limit: because the anchor is the closest whole-genome relative, a sibling that is
itself the closest genome in the collection becomes the anchor and is kept. Curation then
removes any further siblings but not that one; `--exclude-siblings` in the scan (on by
default) is the second line of defence. Check the backbone named in `panel_lineages.tsv`.

`find-references --download ... --curate` curates only what that run downloaded: genomes
already in the download directory are compared against but kept, and appear in
`panel_lineages.tsv` as `sibling-kept` or `redundant-kept` if curation would otherwise
have dropped them. Both need skani (and skDER for dereplication):
`conda install -c bioconda skani skder`.

## Make a collection lineage-ready (`type-lineages`)
Expand Down Expand Up @@ -333,7 +350,7 @@ group); a `parent_group` column in `out/reassortment.tsv` records each segment's
segments are linked only transitively (segment A shares a strain with B, B with C, but A and C share
none) has no single spanning strain, so its `parent_strains` is empty and the text mosaic shows `?`
for it; the verdict is still `clonal` because no pair disagrees. Segments
below `--ani-floor` to every tip, or aligning over too little of their length, or with no resolvable
below `--ani-floor` to every tip, or aligning over less than 50 % of their length, or with no resolvable
dataset, are reported `unassigned` and excluded from the call. Because there is no influenza taxon
alias, flu auto-typing needs the `nextclade` CLI (for `nextclade sort`) or explicit `--dataset
SEGMENT=path` overrides; without either, each flu segment resolves to nothing and is left
Expand Down
Loading
Loading