“A multi-modal transformer-based model for generative visual dialog system” (2025) Applied Computer Science, 21(1), pp. 1–17. doi:10.35784/acs_6856.