Applied Mathematics and Nonlinear Sciences
Journal license

Journal

Applied Mathematics and Nonlinear Sciences


Volume
& Issue

Volume 9, Issue 1


Published
on

November 11, 2024


Pages


DOI

Article

Exploring Semantic Understanding and Generative Modeling in Speech-Text Multimodal Data Fusion

Check for updates


Authors

Haitao Yu Affiliation:
State Grid Tianjin Xintong Company, Tianjin, 300000, China.
, Xuqiang Wang Affiliation:
State Grid Tianjin Electric Power Company, Tianjin, 300000, China.
, Yifan Sun Affiliation:
State Grid Tianjin Xintong Company, Tianjin, 300000, China.
, Yifan Yang Affiliation:
State Grid Tianjin Xintong Company, Tianjin, 300000, China.
and Yan Sun Affiliation:
State Grid Tianjin Xintong Company, Tianjin, 300000, China.


Abstract

Accurate semantic understanding is crucial in the field of human-computer interaction, and it can also greatly improve the comfort of users. In this paper, we use semantic emotion recognition as the research object, collect speech datasets from multiple domains, and extract their semantic features from natural language information. The natural language is digitized using word embedding technology, and then machine learning methods are used to understand the text’s semantics. The attention mechanism is included in the construction of a multimodal Attention-BiLSTM model. The model presented in this paper convergence is achieved in around 20 epochs of training, and the training time and effectiveness are better than those of the other two models. The model in this paper has the highest recognition accuracy. Compared to the S-CBLA model, the recognition accuracy of five semantic emotions, namely happy, angry, sad, sarcastic, and fear, has improved by 24.89%, 15.75%, 1.99%, 2.5%, and 8.5%, respectively. In addition, the probability of correctly recognizing the semantic emotion “Pleasure” in the S-CBLA model is 0.5, while the probability of being recognized as “Angry” is 0.25, which makes it easy to misclassify pleasure as anger. The model in this paper, on the other hand, is capable of distinguishing most semantic emotion types. To conclude, the above experiments confirm the superiority of this paper’s model. This paper’s model improves the accuracy of recognizing semantic emotions and is practical for human-computer interaction.


Keywords

Semantic understanding, Speech dataset, Multimodal data fusion, Attention mechanism, Word embedding, 94A16


Citation

Yu, H., Wang, X., Sun, Y., Yang, Y., & Sun, Y. (2024). Exploring semantic understanding and generative modeling in speech-text multimodal data fusion. Applied Mathematics and Nonlinear Sciences, 9(1). https://doi.org/10.2478/amns-2024-3156

Published by: Engineering Journals

Engineering Journals Logo