Skip to content

Update before release - #9

Merged
matteospanio merged 40 commits into
mainfrom
claude/switch-to-roberta-base-5fq0i
Mar 28, 2026
Merged

matteospanio merged 40 commits into
mainfrom
claude/switch-to-roberta-base-5fq0i

Conversation

@matteospanio

Copy link
Copy Markdown
Member

This pull request focuses on improving documentation clarity, updating configuration defaults, and modernizing training and preprocessing settings for the lilyBERT project. The main changes include a comprehensive rewrite of the README.md with clearer instructions, results, and project structure, updates to configuration files to reflect new model defaults (switching to CodeBERT), and the removal of detailed agent and internal documentation files.

Documentation Improvements:

  • Major rewrite of README.md to include badges, clearer project summary, key results table, installation and reproduction steps, CLI and API references, project structure, and citation information.

Configuration and Defaults Modernization:

  • Updated all model and training configuration defaults to use microsoft/codebert-base instead of bert-base, reduced max_length to 512, increased batch sizes, enabled bf16, and added modern training features such as early stopping, AdamW fused optimizer, and pin memory. [1] [2]
  • Updated SLURM environment configuration to reflect new output directories and disable FSDP by default.
  • Simplified dataset and preprocessing configuration by removing outdated or redundant options (e.g., tokenizer_type, bpe sections) and aligning tokenizer and sharding settings with new defaults. [1] [2] [3] [4] [5] [6]

These changes collectively improve the usability, maintainability, and clarity of the project for both new and existing contributors.

References: [1] [2] [3] [4] [5] [6] [7] [8] [9] [10] [11] [12] [13]

claude and others added 30 commits March 21, 2026 01:09
Replace all BERT-specific imports (BertModel, BertConfig, BertForMaskedLM)
with Auto classes (AutoModel, AutoConfig, AutoModelForMaskedLM) to support
roberta-base and microsoft/codebert-base for both finetuning pretrained
weights and training from scratch. Set type_vocab_size=1 for RoBERTa
compatibility.

https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
Major refactor of the tokenization and training pipeline:

- Add tokenizer_builder.py: loads pretrained tokenizer (roberta-base or
  microsoft/codebert-base) and extends it with ~150 LilyPond musical tokens
  from base_vocabulary() plus [PART_BEGIN]/[PART_END] special tokens.
- Trainer now supports two modes via random_init config flag:
  - random_init=false (default): loads pretrained weights, resizes embeddings
    for the extended tokenizer vocabulary (finetune).
  - random_init=true: initialises architecture from scratch.
- LilyPondMLMDataset converts raw .ly text through the lexer pipeline to
  parser-token representation before feeding to the tokenizer.
- Replace BPE training step in preprocess CLI with tokenizer builder step.
- Update pretokenize/shard workers to always convert through parser tokens.
- Update embed CLI to convert movements through parser tokens.
- Update all config YAML files for the new tokenizer build settings.

https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
Simplify the tokenization pipeline: remove parser-token conversion
(lexer → musical tokens) and feed raw LilyPond text directly to the
pretrained tokenizer.

- tokenizer_builder.py: replace base_vocabulary() semantic tokens with
  ~115 common LilyPond backslash commands (\staccato, \clef, \sustainOn,
  \override, etc.) grouped into articulations, dynamics, musical commands,
  key modes, structural blocks, overrides, and performance/layout.
- Remove [PART_BEGIN]/[PART_END] special tokens.
- Remove all _movement_to_parser_tokens() calls from trainer, dataset,
  pretokenize worker, tokenize worker, and embed CLI.
- Default base model changed from roberta-base to microsoft/codebert-base.

https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
All file-collection helpers now glob for *.ly, *.ily, and *.tely so that
LilyPond include files (.ily) and Texinfo/LilyPond sources (.tely) are
picked up alongside standard .ly scores.

Affected: ly_preprocess, pretokenize, embed, combine, trainer,
preprocessor, repository.

https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
…rride

The script referenced dataset.tokenizer_type (removed in earlier refactor)
and used the old preprocess.bpe.enabled key. Updated to use
preprocess.tokenizer.enabled and removed the --tokenizer-type flag.
The --bpe flag is kept as an alias for --tokenizer for convenience.

https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
Remove text cleaning, variable/score extraction, part/voice separation,
and engraving stripping from the preprocessing pipeline. Files are now
copied as-is to the output directory, with optional data augmentation
preserved (transposition, retrograde, inversion, etc.).

This aligns with the new CodeBERT-based approach where LilyPond files
are treated as plain text (like source code) rather than being split
into movements and cleaned.

- Rewrite preprocessor.py: remove ~1000 lines of parsing/stripping logic
- Simplify CLI config: remove strip and labels_path settings
- Update embed.py: tokenize files directly without movement extraction
- Update scripts/extract_mutopia_embeddings.py for new API
- Update tests for simplified preprocessor interface

https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
CodeBERT's tokenizer defaults to model_max_length=512, causing
"Token indices sequence length is longer than the specified maximum"
warnings when encoding full LilyPond files. Since our windowing code
handles the actual splitting into model-sized chunks, set
model_max_length to a large value in the tokenizer builder and workers.

https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
Update README, CLAUDE.md, and docs/ to articulate the core thesis:
LilyPond is a programming language, so CodeBERT is the natural
pretrained backbone. Fix stale CLI names in CLAUDE.md, add embed
entry point to README, and connect probing/training docs to the
CodeBERT narrative.

https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
Replace _tokenize_mlm_unsharded's in-memory accumulation of all
tokenized samples with streaming writes via ShardWriter. Previously
the entire dataset was held in memory before writing a single .npz,
causing OOM on large corpora even with 512GB RAM. Now each shard is
flushed to disk as it fills (default 8192 samples per shard).

https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
Two root causes for the 512GB OOM during tokenization:

1. ProcessPoolExecutor futures dict held all completed Future objects
   (with their full tokenized results) in memory until the loop ended.
   Fix: del futures[fut] after extracting each result.

2. ShardWriter accumulated movement_ids and base_works for every sample
   in unbounded lists, only used at finalize(). Fix: store metadata
   per-shard in each .npz file instead of accumulating globally.
   ShardedDataset loads metadata from shard files (backward compatible
   with old manifests via per_shard_metadata flag).

https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
- slurm_pipeline.sh: document that both --tokenize and --shard produce
  sharded output with per-shard metadata via streaming ShardWriter
- slurm_train.sh: reference CodeBERT MLM pretraining, clarify examples
- slurm_finetune.sh: note CodeBERT and per-shard metadata format
- extract_emb.sh: reference CodeBERT embeddings

https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
Replace AutoTokenizer with PreTrainedTokenizerFast in MLMPretrainer.
AutoTokenizer.from_pretrained() requires a config.json to auto-detect
the tokenizer class, which doesn't exist in tokenizer-only directories
saved by build_and_save(). PreTrainedTokenizerFast loads directly from
tokenizer.json, which is what the extended CodeBERT tokenizer is.

Fixes: OSError "Repo id must be in the form 'repo_name' or
'namespace/repo_name'" when tokenizer_path is a local directory.

https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
The tokenizer on the cluster is at codebert/tokenizer, not
artifacts/tokenizer.

https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
Codebert accepts sequences up to 512 tokens
Codebert can handle sequences of 512 tokens
Enable mixed precision (bf16), gradient accumulation, fused AdamW,
pinned memory, and early stopping with tuned defaults for faster
continuous pretraining on A40 GPUs (~4x speedup).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
CodeBERT's checkpoint lacks the MLM head (trained with RTD objective),
causing random initialization of lm_head during continuous pretraining.
RoBERTa-base includes the full MLM head and shares the same encoder
architecture. Also removes ignore_mismatched_sizes and disables
ddp_find_unused_parameters to eliminate per-step DDP overhead.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Update slurm training scripts
Remove old hydra flag to specify that is pretraining mode
Make default to false so that we improve training performance
Add configuration parameters to train CLI
Save steps must be a multiple of eval steps
Update batch size 48 -> 64
change FSDP into DDP
do not pin worker processes to specific cpu cores
Use srun to add --cpu-bind=none
Move data into the node
Load shards in RAM for fast data loading
Add tqdm progress bar while loading shards in ram
Read shards as int 32 and print memory load diring dataset ram loading
Use codebert as a default
early stopping triggers too early
Log when early stopping triggers
To restore checkpoints in a reproducible way the possibility to load the trainer configuration from the checkpoint
Increase the amount of slurm memory used in training (we need enough memory to load PDMX in ram).
Change ly-preprocess into preprocess.
Avoid window overlapping and use mean pooling instead of CLS tokens
call model instead of bert attribute in lilybert encoder class
Change tokenizer load function to support codebert-base
Update readme and package metadata before release, adding probing notebook and figure scripts for open science
@matteospanio matteospanio self-assigned this Mar 28, 2026
Fix configuration for smoke test
@matteospanio
matteospanio merged commit 0967c1e into main Mar 28, 2026
3 checks passed
@matteospanio
matteospanio deleted the claude/switch-to-roberta-base-5fq0i branch March 28, 2026 14:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants