Abstract
Communication may greatly be affected by speech and language disorders and therefore, early and proper diagnosis of the condition is the key to effective clinical treatment. The current paper introduces a single artificial intelligence system used in pathological speech classification and improvement through the combination of cross-modal learning and graph-based relational modeling. These two modalities are combined in the proposed system with the help of a cross-modal transformer which identifies the contextual dependencies between the modalities with the help of attention mechanisms. The model uses self-supervised pre training to enhance the quality of representation in a situation where limited labeled data are available and thus it is able to learn both generalized and discriminative speech features. Besides that, a graph neural network is used to model structural dependencies between speech segments, with nodes and edges respectively modeling feature embeddings and temporal continuity and phonetic similarity. This two-sided representation enables the framework to share the analysis of local variations and international speech patterns linked to such disorders like dysarthria and Parkinsonian speech. Moreover, a speech enhancement module is presented that is lightweight and helps to reduce noise interference and to enhance the quality of input signals, thus, boosting the performance of the downstream classification. Both cross-modal transformers and graph neural networks coupled to each other lead to a higher level of robustness, feature discrimination, and interpretability. According to the experimental evidence, the given approach is more effective than the
traditional ones, especially in noisy and low-resource conditions.
Keywords
Speech disorder classification
cross-modal transformer
graph neural networks
self-supervised learning
speech enhancement
multimodal learning
dysarthria
acoustic features
linguistic features
pathological speech analysis
Authors
How to Cite this Article
R. Keerthigadevi, K. Santhanalakshmi (2026).
"CROSS-MODAL TRANSFORMER AND GRAPH NEURAL NETWORK FRAMEWORK FOR ROBUST CLASSIFICATION AND ENHANCEMENT OF PATHOLOGICAL SPEECH".
International Journal of Contemporary Research in Computer Science and Technology,
9(1), pp. 6-11.