MERT
First page

Paper summary
An acoustic music understanding model with large-scale self-supervised training.
Ask this paper
01
Music-specific SSL: Designed specifically for music (not speech/general audio) with appropriate teacher models and training objectives.
02
Multi-teacher design: Combines multiple teacher models to capture different aspects of music (pitch, rhythm, timbre, harmony).
03
Cross-task performance: Outperforms speech and generic audio approaches on music understanding benchmarks (genre, mood, tagging).
04
Music foundation model: Part of the 2023 push toward domain-specific audio foundation models rather than one-size-fits-all speech/audio models.