SheetSage2: Coherent Lead-Sheet Transcription with Synthetic Supervision
Abstract
Transcribing music into a human-readable score requires a coherent understanding of rhythm, harmony, melody, and form. Two obstacles limit this goal: annotated recordings are scarce, and accurate local predictions can still produce inconsistent musical sequences. We present SheetSage2, a unified music transcription framework that combines synthetic data, task-specific structured decoding, and autoregressive distillation. Automatically annotated MIDI, rendered into audio, provides scalable supervision across music understanding tasks. Task-specific structured decoders integrate complementary musical cues and their temporal dependencies to produce musically coherent scores. Autoregressive distillation further retains transcription accuracy without task-specific dynamic programming at inference. Across eight benchmark collections, a single SheetSage2-AR model exceeds the listed prior systems on 12 of 15 benchmark--metric pairs in our evaluation, substantially improving over SheetSage1 and surpassing task-specific models on several benchmarks. Model weights and inference code are publicly available.
Community
We're sharing SheetSage2, a unified framework for turning music recordings into coherent, editable lead sheets and timed musical annotations. It transcribes melody, chords, beats, downbeats, key, and structure.
The framework combines scalable supervision from automatically annotated MIDI rendered into audio, task-specific structured decoding, and autoregressive distillation. A single SheetSage2-AR model outperforms the compared prior systems on 12 of 15 benchmark-metric pairs across eight benchmark collections, while avoiding task-specific dynamic programming at inference.
Model weights and inference code are available, with ABC and MIDI outputs for editing and downstream music workflows.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- StepAudio 3 Music Technical Report (2026)
- Silent Metronome: Rhythmic Grounding for Live Music Accompaniment (2026)
- Note-Level Temporal Grounding of Musical Concepts in Large Audio-Language Models (2026)
- TART: A Modular Tool for Technique-Aware Audio-to-Tablature Guitar Transcription (2026)
- TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models (2026)
- Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning (2026)
- TUTTI: Toward generalizable audio-to-score transcription via fully synthesized data (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.05336 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 2
m-a-p/SheetSage2
Datasets citing this paper 0
No dataset linking this paper