A multi-modal transformer-based model for generative visual dialog system. Applied Computer Science, v. 21, n. 1, p. 1–17, 31 Mar.2025.