1) Y. Fujita, et al.: “End-to-end neural speaker diarization with self-attention,” IEEE ASRU 2019, pp.296-303, 2019.
2) N. Yamashita, et al.: “Improving the naturalness of simulated conversations for end-to-end neural diarization,” The Speaker and Language Recognition Workshop (Odyssey 2022), pp.133-140, 2022.
3) R. Bommasani, et al: “On the opportunities and risks of foundation models,” arXiv:2108.07258 [cs.LG], 2021.
4) N. Makishima, et al.: “Joint autoregressive modeling of end-to-end multi-talker overlapped speech recognition and utterance-level timestamp prediction,” Proc. INTERSPEECH 2023, pp.2913-2917, 2023.
5) A. Baevski, et al.: “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in neural information processing systems, vol.33, pp.12449-12460, 2020.