ATTENTION-ASTFF-NET: A CHANNEL-AWARE ADAPTIVE SPATIOTEMPORAL FUSION NETWORK FOR VEHICLE RECOGNITION

Authors

DOI:

https://doi.org/10.55766/sujst12506

Keywords:

Adaptive Feature Fusion Network, Adaptive Spatiotemporal Fusion Network, Channel-Wise Attention, Vehicle Make Recognition, Vehicle Type Recognition

Abstract

Vehicle recognition systems face challenges in simultaneously capturing spatial appearance characteristics and temporal structural patterns. This study proposes Attention-ASTFF-Net, a channel-aware adaptive spatiotemporal fusion network that enhances vehicle type and vehicle make recognition. The proposed framework extracts spatial features with a convolutional backbone and temporal representations with a convolutional 1D-based module. An adaptive fusion mechanism then combines these features to obtain a comprehensive spatiotemporal representation. Channel-wise attention mechanisms, including Efficient Channel Attention (ECA-Net), Gated Channel Transformation (GCT), Squeeze-and-Excitation (SE) Networks, and Selective Kernel Networks (SKNet), are integrated to emphasize important feature channels and improve discriminative capability. Experimental results on the Vehicle Type Image Dataset Version 2 (VTID2) and the Vehicle Make Image Dataset (VMID) show that the proposed network consistently outperforms the baseline ASTFF-Net and other convolutional neural network architectures. Channel-wise attention integration improves localization, feature selectivity, and prediction stability. The proposed network achieves over 99% accuracy on VTID2 compared with approximately 96% achieved by previous methods, and over 97% on VMID compared with 93.11% achieved by existing approaches. Furthermore, Attention-ASTFF-Net with the SE module demonstrates strong computational efficiency at 5.113 giga floating point operations.

References

Bishop, C. M. (2006). Pattern recognition and machine learning. Springer.

Boonsirisumpun, N., & Surinta, O. (2022a). Fast and accurate deep learning architecture on vehicle type recognition. Current Applied Science and Technology, 22(1). https://doi.org/10.55003/cast.2022.01.22.001

Boonsirisumpun, N., & Surinta, O. (2022b). Ensemble multiple CNNs methods with partial training set for vehicle image classification. Science, Engineering and Health Studies, 16, 22020001. https://doi.org/10.14456/sehs.2022.12

Boonsirisumpun, N., Okafor, E., & Surinta, O. (2024). Vehicle image datasets for image classification. Data in Brief, 53, 110133. https://doi.org/10.1016/j.dib.2024.110133

Cao, Y., Xu, J., Lin, S., Wei, F., & Hu, H. (2019). GCNet: Non-local networks meet squeeze-excitation networks and beyond. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) (pp. 1971-1980). https://doi.org/10.1109/ICCVW.2019.00246

Chen, F., & Tsou, J. Y. (2022). Assessing the effects of convolutional neural network architectural factors on model performance for remote sensing image classification: An in-depth investigation. International Journal of Applied Earth Observation and Geoinformation, 112, 102865. https://doi.org/10.1016/j.jag.2022.102865

Cubuk, E. D., Zoph, B., Mané, D., Vasudevan, V., & Le, Q. V. (2019). AutoAugment: Learning augmentation strategies from data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 113-123). https://doi.org/10.1109/CVPR.2019.00020

Duan, J., Wu, X., Hu, Y., Fu, C., Wang, Z., & He, R. (2023). Iterative embedding distillation for open world vehicle recognition. Pattern Recognition, 135, 109140. https://doi.org/10.1016/j.patcog.2022.109140

Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press.

He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 770-778). https://doi.org/10.1109/CVPR.2016.90

Ho, D., Liang, E., Chen, X., Stoica, I., & Abbeel, P. (2019). Population based augmentation: Efficient learning of augmentation policy schedules. In Proceedings of the 36th International Conference on Machine Learning (pp. 2731-2741).

Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735-1780. https://doi.org/10.1162/neco.1997.9.8.1735

Hu, J., Shen, L., & Sun, G. (2018). Squeeze-and-excitation networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 7132-7141). https://doi.org/10.1109/CVPR.2018.00745

Ji, A., & Ma, X. (2025). Vehicle detection and classification for traffic management and autonomous systems using YOLOv10. Alexandria Engineering Journal, 127, 804-816. https://doi.org/10.1016/j.aej.2025.06.049

Kanagamalliga, S., Kovalan, P., Kiran, K., & Rajalingam, S. (2024). Traffic management through cutting-edge vehicle detection, recognition, and tracking innovations. Procedia Computer Science, 233, 793-800. https://doi.org/10.1016/j.procs.2024.03.268

Ke, X., & Zhang, Y. (2020). Fine-grained vehicle type detection and recognition based on dense attention network. Neurocomputing, 399, 247-257. https://doi.org/10.1016/j.neucom.2020.02.101

Kobiela, D., Groth, J., Hajdasz, M., & Erezman, M. (2024). Vehicle type recognition: A case study of MobileNetV2 for an image classification task. Procedia Computer Science, 246, 3947-3956. https://doi.org/10.1016/j.procs.2024.09.169

Li, X., Wang, W., Hu, X., & Yang, J. (2019). Selective kernel networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 510-519). https://doi.org/10.1109/CVPR.2019.00060

Liu, C., Pu, Z., Li, Y., Jiang, Y., Wang, Y., & Du, Y. (2024). Enabling edge computing ability in view-independent vehicle model recognition. International Journal of Transportation Science and Technology, 14, 73-86. https://doi.org/10.1016/j.ijtst.2023.03.007

Lu, L., Wang, P., & Cao, Y. (2022). A novel part-level feature extraction method for fine-grained vehicle recognition. Pattern Recognition, 131, 108869. https://doi.org/10.1016/j.patcog.2022.108869

Luo, C., Zhu, Y., Jin, L., & Wang, Y. (2020). Learn to augment: Joint data augmentation and network optimization for text recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 13743-13752). https://doi.org/10.1109/CVPR42600.2020.01376

Mehta, R., & Shah, A. (2024). An insight into real time vehicle detection and classification methods using ML/DL based approach. Procedia Computer Science, 235, 598-605. https://doi.org/10.1016/j.procs.2024.04.059

Misra, D., Nalamada, T., Arasanipalai, A. U., & Hou, Q. (2021). Rotate to attend: Convolutional triplet attention module. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV) (pp. 3138-3147). https://doi.org/10.1109/WACV48630.2021.00318

Molchanov, P., Tyree, S., Karras, T., Aila, T., & Kautz, J. (2017). Pruning convolutional neural networks for resource efficient inference. In Proceedings of the International Conference on Learning Representations (ICLR). https://arxiv.org/abs/1611.06440

Ness, S. (2025). Vehicle detection and recognition approach in smart surveillance system: A comparative analysis. Array, 27, 100473. https://doi.org/10.1016/j.array.2025.100473

Phiphitphatphaisit, S., & Surinta, O. (2024). Multi-layer adaptive spatial-temporal feature fusion network for efficient food image recognition. Expert Systems with Applications, 255, 124834. https://doi.org/10.1016/j.eswa.2024.124834

Setthasuravich, P., & Kato, H. (2022). Does the digital divide matter for short-term transportation policy outcomes? A spatial econometric analysis of Thailand. Telematics and Informatics, 72, 101858. https://doi.org/10.1016/j.tele.2022.101858

Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. (2014). Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(1), 1929-1958.

Szegedy, C., Ioffe, S., Vanhoucke, V., & Alemi, A. (2016a). Inception-v4, Inception-ResNet and the impact of residual connections on learning. Proceedings of the AAAI Conference on Artificial Intelligence, 31(1), 4278-4284. https://doi.org/10.1609/aaai.v31i1.11231

Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., & Wojna, Z. (2016b). Rethinking the Inception architecture for computer vision. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 2818–2826). IEEE. https://doi.org/10.1109/CVPR.2016.308

Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning (pp. 6105-6114).

Tan, S. H., Chuah, J. H., Chow, C.-O., Kanesan, J., & Leong, H. Y. (2024). Artificial intelligent systems for vehicle classification: A survey. Engineering Applications of Artificial Intelligence, 129, 107497. https://doi.org/10.1016/j.engappai.2023.107497

Taki, Y., & Zemmouri, E. (2023). Vehicle image classification method using vision transformer. In Artificial Intelligence and Industrial Applications (pp. 221-230). https://doi.org/10.1007/978-3-031-43520-1_19

Wang, Q., Wu, B., Zhu, P., Li, P., Zuo, W., & Hu, Q. (2020). ECA-Net: Efficient channel attention for deep convolutional neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 11531-11539). https://doi.org/10.1109/CVPR42600.2020.01155

Yang, Z., Zhu, L., Wu, Y., & Yang, Y. (2020). Gated channel transformation for visual recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 11791-11800). https://doi.org/10.1109/CVPR42600.2020.01181

Zhang, J., Yang, S., Bo, C., & Zhang, Z. (2021). Vehicle logo detection based on deep convolutional networks. Computers & Electrical Engineering, 90, 107004. https://doi.org/10.1016/j.compeleceng.2021.107004

Zhu, J., Hua, L. S., Chen, B., Kou, W., Rong, J., Zhou, Q., & Sun, Y. (2025). YOLO-DCS: A high-precision edge detection method for latex bowl state and latex trail state. Industrial Crops and Products, 237, 122310. https://doi.org/10.1016/j.indcrop.2025.122310

Downloads

Published

2026-08-01

How to Cite

Jarutan, P., Phiphitphatphaisit, S., Bussabong, Z., Surinta, O., & Rojarath, A. (2026). ATTENTION-ASTFF-NET: A CHANNEL-AWARE ADAPTIVE SPATIOTEMPORAL FUSION NETWORK FOR VEHICLE RECOGNITION. Suranaree Journal of Science and Technology, 33(4), 010440(1–17). https://doi.org/10.55766/sujst12506