Adaptive hybrid neural network architecture with attention for speaker identification in corporate meetings

  • Arseniy D. Potapov, PJSC Sberbank of Russia (Irkutsk, Russia)
  • Nikita D. Lukyanov, Irkutsk National Research Technical University (Irkutsk, Russia)

The relevance of this research stems from the growing need to automate corporate meeting minutes, a labor-intensive and time-consuming process with a high risk of errors. Current speech recognition systems (ASR – Automatic Speech Recognition) perform poorly in multi-user conversations, with overlapping voices, background noise, and the use of specialized vocabulary. Speaker identification (SID – Speaker Identification) is particularly challenging when there are few reference recordings per employee, which is typical in an office environment. Acoustic environments also pose additional challenges: echoing offices, simultaneous speech, and poor voice separation. Therefore, the goal of this study is to develop a hybrid neural network model for accurate transcription and reliable speech attribution in group discussions. This paper analyzes the weaknesses of ASR systems and SID methods and proposes a three-tier architecture: audio stream segmentation (VAD – Voice Activity Detection, SCD – Speaker Change Detection), voice embedding identification with a reference attention mechanism, and speaker-aware transcription. A mathematical justification for the key components is also provided. The theoretical significance lies in the development of an attention-based SID method that effectively solves few-shot learning problems. Unlike traditional approaches that require large labeled corpora, the new architecture learns to match speech fragments with a small number of reference recordings. The effectiveness of the method has been empirically confirmed using real corporate meeting recordings (over 20 hours, five participants). The segmentation module demonstrated a VAD F1 score of 0.885, and the speaker identification module achieved an accuracy of 62.5%, which is three times higher than chance (20%). Transcription was performed using the external Whisper system, which guarantees high recognition accuracy. The results enable the development of a system for automatic meeting minutes, and the architecture supports a large number of participants and is easily adapted to various communication formats: from negotiations to client meetings. The functionality of the hybrid neural network model can be used to create various tools for corporate meeting minutes 

Automatic speech recognition, speaker identification, hybrid neural network architecture, attention mechanism, voice embeddings, audio stream segmentation, transcription, corporate meetings

2026-09-03

Copyright (c) 2026 Information and mathematical technologies in science and management
Back