Repository navigation
Update before release - #9
Merged
Merged
Conversation
Replace all BERT-specific imports (BertModel, BertConfig, BertForMaskedLM) with Auto classes (AutoModel, AutoConfig, AutoModelForMaskedLM) to support roberta-base and microsoft/codebert-base for both finetuning pretrained weights and training from scratch. Set type_vocab_size=1 for RoBERTa compatibility. https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
Major refactor of the tokenization and training pipeline:
- Add tokenizer_builder.py: loads pretrained tokenizer (roberta-base or
microsoft/codebert-base) and extends it with ~150 LilyPond musical tokens
from base_vocabulary() plus [PART_BEGIN]/[PART_END] special tokens.
- Trainer now supports two modes via random_init config flag:
- random_init=false (default): loads pretrained weights, resizes embeddings
for the extended tokenizer vocabulary (finetune).
- random_init=true: initialises architecture from scratch.
- LilyPondMLMDataset converts raw .ly text through the lexer pipeline to
parser-token representation before feeding to the tokenizer.
- Replace BPE training step in preprocess CLI with tokenizer builder step.
- Update pretokenize/shard workers to always convert through parser tokens.
- Update embed CLI to convert movements through parser tokens.
- Update all config YAML files for the new tokenizer build settings.
https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
Simplify the tokenization pipeline: remove parser-token conversion (lexer → musical tokens) and feed raw LilyPond text directly to the pretrained tokenizer. - tokenizer_builder.py: replace base_vocabulary() semantic tokens with ~115 common LilyPond backslash commands (\staccato, \clef, \sustainOn, \override, etc.) grouped into articulations, dynamics, musical commands, key modes, structural blocks, overrides, and performance/layout. - Remove [PART_BEGIN]/[PART_END] special tokens. - Remove all _movement_to_parser_tokens() calls from trainer, dataset, pretokenize worker, tokenize worker, and embed CLI. - Default base model changed from roberta-base to microsoft/codebert-base. https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
All file-collection helpers now glob for *.ly, *.ily, and *.tely so that LilyPond include files (.ily) and Texinfo/LilyPond sources (.tely) are picked up alongside standard .ly scores. Affected: ly_preprocess, pretokenize, embed, combine, trainer, preprocessor, repository. https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
…rride The script referenced dataset.tokenizer_type (removed in earlier refactor) and used the old preprocess.bpe.enabled key. Updated to use preprocess.tokenizer.enabled and removed the --tokenizer-type flag. The --bpe flag is kept as an alias for --tokenizer for convenience. https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
Remove text cleaning, variable/score extraction, part/voice separation, and engraving stripping from the preprocessing pipeline. Files are now copied as-is to the output directory, with optional data augmentation preserved (transposition, retrograde, inversion, etc.). This aligns with the new CodeBERT-based approach where LilyPond files are treated as plain text (like source code) rather than being split into movements and cleaned. - Rewrite preprocessor.py: remove ~1000 lines of parsing/stripping logic - Simplify CLI config: remove strip and labels_path settings - Update embed.py: tokenize files directly without movement extraction - Update scripts/extract_mutopia_embeddings.py for new API - Update tests for simplified preprocessor interface https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
CodeBERT's tokenizer defaults to model_max_length=512, causing "Token indices sequence length is longer than the specified maximum" warnings when encoding full LilyPond files. Since our windowing code handles the actual splitting into model-sized chunks, set model_max_length to a large value in the tokenizer builder and workers. https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
Update README, CLAUDE.md, and docs/ to articulate the core thesis: LilyPond is a programming language, so CodeBERT is the natural pretrained backbone. Fix stale CLI names in CLAUDE.md, add embed entry point to README, and connect probing/training docs to the CodeBERT narrative. https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
Replace _tokenize_mlm_unsharded's in-memory accumulation of all tokenized samples with streaming writes via ShardWriter. Previously the entire dataset was held in memory before writing a single .npz, causing OOM on large corpora even with 512GB RAM. Now each shard is flushed to disk as it fills (default 8192 samples per shard). https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
Two root causes for the 512GB OOM during tokenization: 1. ProcessPoolExecutor futures dict held all completed Future objects (with their full tokenized results) in memory until the loop ended. Fix: del futures[fut] after extracting each result. 2. ShardWriter accumulated movement_ids and base_works for every sample in unbounded lists, only used at finalize(). Fix: store metadata per-shard in each .npz file instead of accumulating globally. ShardedDataset loads metadata from shard files (backward compatible with old manifests via per_shard_metadata flag). https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
- slurm_pipeline.sh: document that both --tokenize and --shard produce sharded output with per-shard metadata via streaming ShardWriter - slurm_train.sh: reference CodeBERT MLM pretraining, clarify examples - slurm_finetune.sh: note CodeBERT and per-shard metadata format - extract_emb.sh: reference CodeBERT embeddings https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
Replace AutoTokenizer with PreTrainedTokenizerFast in MLMPretrainer. AutoTokenizer.from_pretrained() requires a config.json to auto-detect the tokenizer class, which doesn't exist in tokenizer-only directories saved by build_and_save(). PreTrainedTokenizerFast loads directly from tokenizer.json, which is what the extended CodeBERT tokenizer is. Fixes: OSError "Repo id must be in the form 'repo_name' or 'namespace/repo_name'" when tokenizer_path is a local directory. https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
The tokenizer on the cluster is at codebert/tokenizer, not artifacts/tokenizer. https://claude.ai/code/session_01N8VoY5ozyNWUBUxzmH7p3Z
Codebert accepts sequences up to 512 tokens
Codebert can handle sequences of 512 tokens
Enable mixed precision (bf16), gradient accumulation, fused AdamW, pinned memory, and early stopping with tuned defaults for faster continuous pretraining on A40 GPUs (~4x speedup). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
CodeBERT's checkpoint lacks the MLM head (trained with RTD objective), causing random initialization of lm_head during continuous pretraining. RoBERTa-base includes the full MLM head and shares the same encoder architecture. Also removes ignore_mismatched_sizes and disables ddp_find_unused_parameters to eliminate per-step DDP overhead. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Update slurm training scripts
Remove old hydra flag to specify that is pretraining mode
Make default to false so that we improve training performance
Add configuration parameters to train CLI
Save steps must be a multiple of eval steps
Update batch size 48 -> 64
change FSDP into DDP
do not pin worker processes to specific cpu cores
Use srun to add --cpu-bind=none
Move data into the node
Load shards in RAM for fast data loading
Add tqdm progress bar while loading shards in ram
Read shards as int 32 and print memory load diring dataset ram loading
Use codebert as a default
early stopping triggers too early
Log when early stopping triggers
To restore checkpoints in a reproducible way the possibility to load the trainer configuration from the checkpoint
Increase the amount of slurm memory used in training (we need enough memory to load PDMX in ram). Change ly-preprocess into preprocess.
Avoid window overlapping and use mean pooling instead of CLS tokens
call model instead of bert attribute in lilybert encoder class
Change tokenizer load function to support codebert-base
Update readme and package metadata before release, adding probing notebook and figure scripts for open science
Fix configuration for smoke test
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This pull request focuses on improving documentation clarity, updating configuration defaults, and modernizing training and preprocessing settings for the lilyBERT project. The main changes include a comprehensive rewrite of the
README.mdwith clearer instructions, results, and project structure, updates to configuration files to reflect new model defaults (switching to CodeBERT), and the removal of detailed agent and internal documentation files.Documentation Improvements:
README.mdto include badges, clearer project summary, key results table, installation and reproduction steps, CLI and API references, project structure, and citation information.Configuration and Defaults Modernization:
microsoft/codebert-baseinstead ofbert-base, reducedmax_lengthto 512, increased batch sizes, enabled bf16, and added modern training features such as early stopping, AdamW fused optimizer, and pin memory. [1] [2]tokenizer_type,bpesections) and aligning tokenizer and sharding settings with new defaults. [1] [2] [3] [4] [5] [6]These changes collectively improve the usability, maintainability, and clarity of the project for both new and existing contributors.
References: [1] [2] [3] [4] [5] [6] [7] [8] [9] [10] [11] [12] [13]