Cross-Architecture Evaluation of Deep Learning Models for Multimodal Biometric Verification with Score-Level Fusion

Authors

  • Rowida Jasim Alazawi Technical College of Management Middle Technical University Baghdad, Iraq Author https://orcid.org/0009-0001-4890-3431
  • Nada Jasim Habeeb Technical College of Management, Middle Technical University, Baghdad, Iraq. Author
  • Alaa Jabbar Qasim Almaliki School of Computing, Universiti Utara Malaysia, 06010 Sintok, Kedah Darul Aman, Malaysia. Author

DOI:

https://doi.org/10.24017/science.2026.2.2

Keywords:

Biometric Verification, Deep Learning, Multimodal Fusion, Score-Level Fusion, CNN, Transformer

Abstract

Comparing the deep learning architectures used for biometric verification in a fair manner has remained a difficult task. Researchers typically rely on different datasets, employ varying preprocessing steps, and adopt inconsistent evaluation criteria, all of which muddy the picture when trying to decide whether one network truly outperforms another. This work attempts to cut through that noise by fixing every variable except the backbone architecture itself. Residual Network-50) ResNet50), Efficient Network Version 2 Small (EfficientNetV2-S), and the Swin Transformer-Tiny (Swin-T) were trained and tested on the same LUTBIO and XJTU datasets, applying identical subject-disjoint splits across face, fingerprint, and palmprint modalities. The analysis measures Equal Error Rate (EER), Area Under the Curve, and True Acceptance Rate at False Acceptance Rate = 0.1%, and further examines five distinct score-level fusion strategies across every possible two- and three-modality pairing. The single-modality experiments reveal a clear pattern: Swin-T achieves a perfect 0.000% EER on faces, ResNet50 handles fingerprints best at 1.957% EER, and EfficientNetV2-S leads on palmprints with 0.140% EER. On LUTBIO, 10 out of 12 tested fusion configurations reach a reported EER of 0.000%. This study reports the effect sizes and selected confidence intervals for the key comparisons, allowing both their statistical and practical significance to be assessed. The results show that no single architecture is universally superior; the best-performing backbone is modality-dependent. Score-level fusion consistently improved verification accuracy over the best single modality, substantially reducing error rates. These findings imply that architecture choices for biometric systems should be modality-aware and that score-level fusion is an appropriate method for robust multimodal verification. The controlled protocol additionally provides a reproducible framework for fair architectural comparison in future biometric research.

References

[1] S. Minaee, A. Abdolrashidi, H. Su, M. Bennamoun, and D. Zhang, “Biometrics recognition using deep learning: a survey,” Artificial Intelligence Review, vol. 56, no. 8, pp. 8647–8695, Aug. 2023, doi: 10.1007/s10462-022-10237-x. DOI: https://doi.org/10.1007/s10462-022-10237-x

[2] M. Rodrigo, C. Cuevas, and N. García, “Comprehensive comparison between vision transformers and convolutional neural networks for face recognition tasks,” Scientific Reports, vol. 14, no. 1, p. 21392, Sep. 2024, doi: 10.1038/s41598-024-72254-w. DOI: https://doi.org/10.1038/s41598-024-72254-w

[3] K. Han et al., “A survey on vision transformer,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 1, pp. 87–110, Jan. 2023, doi: 10.1109/TPAMI.2022.3152247. DOI: https://doi.org/10.1109/TPAMI.2022.3152247

[4] J. Maurício, I. Domingues, and J. Bernardino, “Comparing vision transformers and convolutional neural networks for image classification: a literature review,” Applied Sciences, vol. 13, no. 9, p. 5521, Apr. 2023, doi: 10.3390/app13095521. DOI: https://doi.org/10.3390/app13095521

[5] R. Garg, G. Singh, A. Singh, and M. P. Singh, “Fingerprint recognition using convolution neural network with inver-sion and augmented techniques,” Systems and Soft Computing, vol. 6, p. 200106, Dec. 2024, doi: 10.1016/j.sasc.2024.200106. DOI: https://doi.org/10.1016/j.sasc.2024.200106

[6] C. Gao, Z. Yang, W. Jia, L. Leng, B. Zhang, and A. B. J. Teoh, “Deep learning in palmprint recognition: a comprehen-sive survey,” IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 56, no. 3, pp. 2143–2162, Mar. 2026, doi: 10.1109/TSMC.2025.3649416. DOI: https://doi.org/10.1109/TSMC.2025.3649416

[7] A. A. Ross, A. K. Jain, and K. Nandakumar, Handbook of Multibiometrics, vol. 6, International Series on Biometrics. Boston, MA: Kluwer Academic Publishers, May. 2006. doi: 10.1007/0-387-33123-9. DOI: https://doi.org/10.1007/0-387-33123-9

[8] R. Yang, Q. Zhang, L. Meng, C. Wang, and Y. Hu, LUTBIO multimodal biometric database. Mendeley data, Jan. 27, 2025. Accessed: Jun. 12, 2026. [Online]. doi: 10.17632/jszw485f8j.6.

[9] ISO/IEC 19795-1:2021 Information Technology—Biometric Performance Testing and Reporting—Part 1: Principles and Framework. Geneva, Switzerland: International Organization for Standardization (ISO), 2021. Accessed: Jun. 12, 2026. [Online]. Available: https://www.iso.org/standard/73515.html.

[10] D. Shang, X. Zhang, J. Han, and X. Xu, “MultiModal-database-XJTU: An available database for biometrics recogni-tion with its performance testing,” in 2017 IEEE 3rd Information Technology and Mechatronics Engineering Conference (ITOEC), Chongqing, China, Oct. 2017, pp. 521–526, doi: 10.1109/ITOEC.2017.8122351. DOI: https://doi.org/10.1109/ITOEC.2017.8122351

[11] M. Wang and W. Deng, “Deep face recognition: a survey,” Neurocomputing, vol. 429, pp. 215–244, Mar. 2021, doi: 10.1016/j.neucom.2020.10.081. DOI: https://doi.org/10.1016/j.neucom.2020.10.081

[12] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in 2016 IEEE Conference on Com-puter Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, Jun. 2016, pp. 770–778, doi: 10.1109/CVPR.2016.90. DOI: https://doi.org/10.1109/CVPR.2016.90

[13] M. Tan and Q. Le, “EfficientNetV2: smaller models and faster training,” in Proceedings of the 38th International Con-ference on Machine Learning (ICML), Proceedings of Machine Learning Research (PMLR), vol. 139, Jul. 2021, pp. 10096–10106. Accessed: Jun. 13, 2026. [Online]. Available: https://proceedings.mlr.press/v139/tan21a.html.

[14] Z. Liu et al., “Swin Transformer: hierarchical vision transformer using shifted windows,” in 2021 IEEE/CVF Interna-tional Conference on Computer Vision (ICCV), Montreal, QC, Canada, Oct. 2021, pp. 9992–10002, doi: 10.1109/ICCV48922.2021.00986. DOI: https://doi.org/10.1109/ICCV48922.2021.00986

[15] N. Peri, J. Gleason, C. D. Castillo, T. Bourlai, V. M. Patel, and R. Chellappa, “A synthesis-based approach for ther-mal-to-visible face verification,” in 2021 16th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2021), Jodhpur, India, Dec. 2021, pp. 1–8, doi: 10.1109/FG52635.2021.9666943. DOI: https://doi.org/10.1109/FG52635.2021.9666943

[16] Z. He, J. Zhang, L. Pang, and E. Liu, “PFVNet: a partial fingerprint verification network learned from large finger-print matching,” IEEE Transactions on Information Forensics and Security, vol. 17, pp. 3706–3719, Sep. 2022, doi: 10.1109/TIFS.2022.3209869. DOI: https://doi.org/10.1109/TIFS.2022.3209869

[17] D. Brown and K. Bradshaw, “Deep palmprint recognition with alignment and augmentation of limited training samples,” SN Computer Science, vol. 3, no. 1, p. 11, Jan. 2022, doi: 10.1007/s42979-021-00859-3. DOI: https://doi.org/10.1007/s42979-021-00859-3

[18] W. Su, Y. Wang, K. Li, P. Gao, and Y. Qiao, “Hybrid token transformer for deep face recognition,” Pattern Recogni-tion, vol. 139, p. 109443, Jul. 2023, doi: 10.1016/j.patcog.2023.109443. DOI: https://doi.org/10.1016/j.patcog.2023.109443

[19] A. Hattab and A. Behloul, “Face-Iris multimodal biometric recognition system based on deep learning,” Multimedia Tools and Applications, vol. 83, no. 14, pp. 43349–43376, Oct. 2023, doi: 10.1007/s11042-023-17337-y. DOI: https://doi.org/10.1007/s11042-023-17337-y

[20] Y. Yang, L. Fei, A. H. Alshehri, S. Zhao, W. Sun, and S. Teng, “Joint multi-type feature learning for multi-modality FKP recognition,” Engineering Applications of Artificial Intelligence, vol. 126, p. 106960, Nov. 2023, doi: 10.1016/j.engappai.2023.106960. DOI: https://doi.org/10.1016/j.engappai.2023.106960

[21] L. Su, L. Fei, B. Zhang, S. Zhao, J. Wen, and Y. Xu, “Complete region of interest for unconstrained palmprint recog-nition,” IEEE Transactions on Image Processing, vol. 33, pp. 3662–3675, Jun. 2024, doi: 10.1109/TIP.2024.3407666. DOI: https://doi.org/10.1109/TIP.2024.3407666

[22] Y. Wu and J. Hu, “Deep and shallow feature fusion in feature score level for palmprint recognition,” IET Biometrics, vol. 2024, no. 1, p. 5683547, Jan. 2024, doi: 10.1049/2024/5683547. DOI: https://doi.org/10.1049/2024/5683547

[23] C. Lin et al., “An unconstrained palmprint region of interest extraction method based on lightweight networks,” PLoS ONE, vol. 19, no. 8, p. e0307822, Aug. 2024, doi: 10.1371/journal.pone.0307822. DOI: https://doi.org/10.1371/journal.pone.0307822

[24] D. Fan, X. Liang, W. Jia, J. Chen, and D. Zhang, “A novel hybrid fusion combining palmprint and palm vein for large-scale palm-based recognition,” IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 54, no. 7, pp. 4471–4484, Jul. 2024, doi: 10.1109/TSMC.2024.3382877. DOI: https://doi.org/10.1109/TSMC.2024.3382877

[25] S. A. Grosz, A. Godbole, and A. K. Jain, “Mobile contactless palmprint recognition: use of multiscale, multimodel embeddings,” IEEE Transactions on Information Forensics and Security, vol. 19, pp. 8428–8440, Jun. 2024, doi: 10.1109/TIFS.2024.3413631. DOI: https://doi.org/10.1109/TIFS.2024.3413631

[26] A. Hossain, R. Sumi, and S. Schuckers, “Evaluating deep learning-based face recognition for infants and toddlers: impact of age across developmental stages,” in 2025 IEEE International Joint Conference on Biometrics (IJCB), Osaka, Japan, Sep. 2025, pp. 1–9, doi: 10.1109/IJCB65343.2025.11410705. DOI: https://doi.org/10.1109/IJCB65343.2025.11410705

[27] Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, pp. 436–444, May 2015, doi: 10.1038/nature14539. DOI: https://doi.org/10.1038/nature14539

[28] X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer, “Scaling vision transformers,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, Jun. 2022, pp. 1204–1213, doi: 10.1109/CVPR52688.2022.01179. DOI: https://doi.org/10.1109/CVPR52688.2022.01179

[29] A. Ross and A. Jain, “Information fusion in biometrics,” Pattern Recognition Letters, vol. 24, no. 13, pp. 2115–2125, Sep. 2003, doi: 10.1016/S0167-8655(03)00079-5. DOI: https://doi.org/10.1016/S0167-8655(03)00079-5

[30] Y. Feng and A. Kumar, “BEST: Building evidences from scattered templates for accurate contactless palmprint recognition,” Pattern Recognition, vol. 138, p. 109422, Jun. 2023, doi: 10.1016/j.patcog.2023.109422. DOI: https://doi.org/10.1016/j.patcog.2023.109422

[31] M. Khatri and A. Sharma, “Deep learning approach based on iris, face, and palmprint fusion for multimodal bio-metric recognition system,” International Journal of Performability Engineering, vol. 19, no. 6, p. 407, Jun. 2023, doi: 10.23940/ijpe.23.06. p6.407416. DOI: https://doi.org/10.23940/ijpe.23.06.p6.407416

[32] S. A. El_Rahman and A. S. Alluhaidan, “Enhanced multimodal biometric recognition systems based on deep learn-ing and traditional methods in smart environments,” PLoS ONE, vol. 19, no. 2, p. e0291084, Feb. 2024, doi: 10.1371/journal.pone.0291084. DOI: https://doi.org/10.1371/journal.pone.0291084

[33] S. Li, L. Fei, B. Zhang, X. Ning, and L. Wu, “Hand-based multimodal biometric fusion: a review,” Information Fu-sion, vol. 109, p. 102418, Sep. 2024, doi: 10.1016/j.inffus.2024.102418. DOI: https://doi.org/10.1016/j.inffus.2024.102418

[34] S. Dargan and M. Kumar, “A comprehensive survey on the biometric recognition systems based on physiological and behavioral modalities,” Expert Systems with Applications, vol. 143, p. 113114, Apr. 2020, doi: 10.1016/j.eswa.2019.113114. DOI: https://doi.org/10.1016/j.eswa.2019.113114

[35] ISO/IEC 24745:2022 Information Security, Cybersecurity and Privacy Protection—Biometric Information Protection. Ge-neva, Switzerland: International Organization for Standardization (ISO), Feb. 2022. Accessed: Jun. 12, 2026. [Online]. Available: https://www.iso.org/standard/75302.html.

[36] K. Zhang, Z. Zhang, Z. Li, and Y. Qiao, “Joint face detection and alignment using multitask cascaded convolutional networks,” IEEE Signal Processing Letters, vol. 23, no. 10, pp. 1499–1503, Oct. 2016, doi: 10.1109/LSP.2016.2603342. DOI: https://doi.org/10.1109/LSP.2016.2603342

[37] K. Zuiderveld, “Contrast limited adaptive histogram equalization,” in Graphics Gems IV, P. S. Heckbert, Ed. San Diego, CA: Academic Press, Aug. 1994, pp. 474–485. DOI: https://doi.org/10.1016/B978-0-12-336156-1.50061-6

[38] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: a large-scale hierarchical image database,” in 2009 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Miami, FL, USA, Jun. 2009, pp. 248–255, doi: 10.1109/CVPR.2009.5206848. DOI: https://doi.org/10.1109/CVPR.2009.5206848

[39] J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “ArcFace: additive angular margin loss for deep face recognition,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, Jun. 2019, pp. 4685–4694, doi: 10.1109/CVPR.2019.00482. DOI: https://doi.org/10.1109/CVPR.2019.00482

[40] H. Wang et al., “CosFace: large margin cosine loss for deep face recognition,” in 2018 IEEE/CVF Conference on Com-puter Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, Jun. 2018, pp. 5265–5274, doi: 10.1109/CVPR.2018.00552. DOI: https://doi.org/10.1109/CVPR.2018.00552

[41] F. Wang, J. Cheng, W. Liu, and H. Liu, “Additive margin softmax for face verification,” IEEE Signal Processing Let-ters, vol. 25, no. 7, pp. 926–930, Jul. 2018, doi: 10.1109/LSP.2018.2822810. DOI: https://doi.org/10.1109/LSP.2018.2822810

Downloads

Published

21-07-2026

Issue

Section

Pure and Applied Science

Similar Articles

91-100 of 123

You may also start an advanced similarity search for this article.