Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Music genre classification models generally achieve good performance on their own training sets, but fail to perform well on sets of varying quality, production style, artist composition, genre taxonomy and annotation practice. A dual-branch architecture is proposed in this paper, MERT-SpecAMGCNet, that combines the pretrained MERT waveform representations and a frequency-aware log-Mel spectral branch. The spectral branch uses multi-scale gated convolutions, chan- nel recalibration and low-, mid-, high-frequency region fusion to capture local time–frequency evidence while gated cross-representation attention selectively fuses the spectral tokens with the temporal embeddings from MERT. This fused representation is further enhanced by temporal modelling using Conformer and multi-head attentive statistics pooling. The model is evaluated on the GTZAN 10-class classification, GTZAN-to-FMA Shared-3 zero-shot transfer, full target- domain fine-tuning (Shared-3) and FMA Small target-domain fine-tuning (8-class supervised classification). MERT-SpecAMGCNet achieved 89.60±0.55% accu- racy and 89.90 ± 0.60% macro-F1 on GTZAN, 65.00 ± 0.71% accuracy and 62.77 ± 0.82% macro-F1 in zero-shot GTZAN-to-FMA Shared-3 transfer, 89.00 ± 0.71% accuracy and 88.75 ± 0.74% macro-F1 after full FMA Shared- 3 fine-tuning, and 81.60 ± 0.76% accuracy and 81.44 ± 0.84% macro-F1 on FMA Small. Baseline and ablation studies demonstrate that each of the four components (frequency-aware spectral modelling, gated fusion, temporal refine- ment, attentive pooling) makes a contribution to the performance. The results show that explicit spectral evidence can be applied along with the pre-trained music embeddings for the transfer learning-based music genre classification.</p>

Show More

Keywords

spectral classification macrof1 music genre

Related Articles

PORE

About

Connect