Adaptive Transformer Learning for Robust Intent Classification and Slot Filling

Authors

  • Saeed Biabani Mahalli Department of Electrical and Computer Engineering University of Science and Technology of Mazandaran, Iran.
  • Nastaran Zarafshan Department of Computer Engineering Shahrood University of Technology Semnan, Iran.
  • Jamshid Pirgazi * Department of Electrical and Computer Engineering University of Science and Technology of Mazandaran, Iran. https://orcid.org/0000-0002-2461-1143

https://doi.org/10.48314/jidcm.v2i2.91

Abstract

This study proposes a lightweight transformer-based model for intent classification and slot filling in conversational systems. By combining Multi-Head Attention (MHA), Root Mean Square (RMS) normalization, Swish Gated Linear Unit (SwiGLU) feedforward layers, and two-stage adaptive training on hard samples, the model improves robustness, handles imbalanced data, and achieves stronger accuracy and generalization in Natural Language Understanding (NLU) tasks.

Keywords:

Intent classification, Slot filling, Transformer network, Multi-head attention, Imbalanced data, Deep learning

References

  1. [1] Tur, G., & De Mori, R. (2011). Spoken language understanding: Systems for extracting semantic information from speech. John Wiley & Sons. https://doi.org/10.1002/9781119992691

  2. [2] Ray, S., & Craven, M. (2005). Supervised versus multiple instance learning: An empirical comparison. Proceedings of the 22nd international conference on machine learning (pp. 697–704). Association for Computing Machinery (ACM). https://doi.org/10.1145/1102351.1102439

  3. [3] Pieraccini, R., & Huerta, J. (2005). Where do we go from here? Research and commercial spoken dialog systems. Proceedings of the 6th SIGdial workshop on discourse and dialogue (pp. 1-10). Special Interest Group on Discourse and Dialogue (SIGdial). https://doi.org/10.18653/v1/2005.sigdial-1.1

  4. [4] Yao, K., Zweig, G., Hwang, M. Y., Shi, Y., & Yu, D. (2013). Recurrent neural networks for language understanding. IEEE transactions on audio, speech, and language processing, 21(10), 2093–2103. https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/338_Paper.pdf

  5. [5] Goo, C. W., Gao, G., Hsu, Y. K., Huo, C. L., Chen, T. C., Hsu, K. W., & Chen, Y. N. (2018). Slot gated modeling for joint slot filling and intent prediction. Proceedings of the 2018 conference of the north american chapter of the association for computational linguistics: Human language technologies, volume 2 (short papers) (pp. 753–757). Association for Computational Linguistics. https://doi.org/10.18653/v1/N18-2118

  6. [6] Zhang, X., & Wang, H. (2016). A joint model of intent determination and slot filling for spoken language understanding. Proceedings of the twenty-fifth international joint conference on artificial intelligence (pp. 2993–2999). AAAI Press. https://zxdcs.github.io/pdf/spoken_language_understanding.pdf

  7. [7] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30. https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf

  8. [8] Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). Bert: Pre training of deep bidirectional transformers for language understanding. Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: Human language technologies, volume 1 (long and short papers) (pp. 4171-4186). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/N19-1423

  9. [9] Radford, A., Narasimhan, K., Salimans, T., Sutskever, I. (2018). Improving language understanding by generative pretraining. San Francisco, CA, USA. https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf

  10. [10] Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., ... & Stoyanov, V. (2019). Roberta: A robustly optimized bert pretraining approach. https://doi.org/10.48550/arXiv.1907.11692

  11. [11] Chen, Q., Zhuo, Z., & Wang, W. (2019). Bert for joint intent classification and slot filling. https://doi.org/10.48550/arXiv.1902.10909

  12. [12] Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R. R., & Le, Q. V. (2019). Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems, 32. Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2019/file/dc6a7e655d7e5840e66733e9ee67cc69-Paper.pdf

  13. [13] Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., & Liu, P. J. (2020). Exploring the limits of transfer learning with a unified text to text transformer. Journal of machine learning research, 21(140), 1–67. https://www.jmlr.org/papers/volume21/20-074/20-074.pdf

  14. [14] Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., ... & Stoyanov, V. (2020). Unsupervised cross lingual representation learning at scale. Proceedings of the 58th annual meeting of the association for computational linguistics (pp. 8440-8451). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2020.acl-main.747

  15. [15] Castellucci, G., Bellomaria, V., Favalli, A., & Romagnoli, R. (2019). Multi lingual intent detection and slot filling in a joint bert based model. https://doi.org/10.48550/arxiv.1907.02884

  16. [16] Jiang, H., He, P., Chen, W., Liu, X., Gao, J., & Zhao, T. (2020). Smart: Robust and efficient fine tuning for pre trained natural language models through principled regularized optimization. Proceedings of the 58th annual meeting of the association for computational linguistics (pp. 2177-2190). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.197

  17. [17] Ganesh, P., Chen, Y., Lou, X., Khan, M. A., Yang, Y., Sajjad, H., Nakov, P., Chen, D., & Winslett, M. (2021). Compressing larges cale transformer based models: A case study on bert. Transactions of the association for computational linguistics, 9, 1061–1080. https://doi.org/10.1162/tacl_a_00413

  18. [18] Gu, S., Zhang, J., Meng, F., Feng, Y., Xie, W., Zhou, J., & Yu, D. (2020). Token level adaptive training for neural machine translation. Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP) (pp. 1035–1046). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-main.76

  19. [19] Shazeer, N. (2020). Glu variants improve transformer. https://doi.org/10.48550/arXiv.2002.05202

Published

2026-06-17

How to Cite

Biabani Mahalli, S. ., Zarafshan, N. ., & Pirgazi, J. . (2026). Adaptive Transformer Learning for Robust Intent Classification and Slot Filling. Journal of Intelligent Decision and Computational Modelling, 2(2), 137-154. https://doi.org/10.48314/jidcm.v2i2.91

Similar Articles

1-10 of 19

You may also start an advanced similarity search for this article.