VISCOUNTH: A Large-Scale Multilingual Visual Question Answering Dataset for Cultural Heritage

IRIS

Visual question answering has recently been settled as a fundamental multi-modal reasoning task of artificial intelligence that allows users to get information about visual content by asking questions in natural language. In the cultural heritage domain this task can contribute to assist visitors in museums and cultural sites, thus increasing engagement. However, the development of visual question answering models for cultural heritage is prevented by the lack of suitable large-scale datasets. To meet this demand, we built a large-scale heterogeneous and multilingual (Italian and English) dataset for cultural heritage that comprises approximately 500K Italian cultural assets and 6.5M question-answer pairs. We propose a novel formulation of the task that requires reasoning over both the visual content and an associated natural language description, and present baselines for this task. Results show that the current state of the art is reasonably effective, but still far from satisfactory, therefore further research is this area is recommended. Nonetheless, we also present a holistic baseline to address visual and contextual questions and foster future research on the topic.

Becattini, F., Bongini, P., Bulla, L., Del Bimbo, A., Marinucci, L., Mongiovì, M., et al. (2023). VISCOUNTH: A Large-Scale Multilingual Visual Question Answering Dataset for Cultural Heritage. ACM TRANSACTIONS ON MULTIMEDIA COMPUTING, COMMUNICATIONS AND APPLICATIONS [10.1145/3590773].

VISCOUNTH: A Large-Scale Multilingual Visual Question Answering Dataset for Cultural Heritage

Becattini, Federico;Bongini, Pietro;Bulla, Luana;Del Bimbo, Alberto;Marinucci, Ludovica;Mongiovì, Misael;Presutti, Valentina

2023-01-01

Abstract

Visual question answering has recently been settled as a fundamental multi-modal reasoning task of artificial intelligence that allows users to get information about visual content by asking questions in natural language. In the cultural heritage domain this task can contribute to assist visitors in museums and cultural sites, thus increasing engagement. However, the development of visual question answering models for cultural heritage is prevented by the lack of suitable large-scale datasets. To meet this demand, we built a large-scale heterogeneous and multilingual (Italian and English) dataset for cultural heritage that comprises approximately 500K Italian cultural assets and 6.5M question-answer pairs. We propose a novel formulation of the task that requires reasoning over both the visual content and an associated natural language description, and present baselines for this task. Results show that the current state of the art is reasonably effective, but still far from satisfactory, therefore further research is this area is recommended. Nonetheless, we also present a holistic baseline to address visual and contextual questions and foster future research on the topic.

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno
	
				2023
			
	Rivista su cui è pubblicata l'opera
	
				ACM TRANSACTIONS ON MULTIMEDIA COMPUTING, COMMUNICATIONS AND APPLICATIONS
			
	Citazione
	
				Becattini, F., Bongini, P., Bulla, L., Del Bimbo, A., Marinucci, L., Mongiovì, M., et al. (2023). VISCOUNTH: A Large-Scale Multilingual Visual Question Answering Dataset for Cultural Heritage. ACM TRANSACTIONS ON MULTIMEDIA COMPUTING, COMMUNICATIONS AND APPLICATIONS [10.1145/3590773].
			
	Appare nelle tipologie:
	
				1.1 Articolo in rivista

File in questo prodotto:

File	Dimensione	Formato
3590773.pdf accesso aperto Tipologia: Post-print Licenza: PUBBLICO - Pubblico con Copyright Dimensione 1.23 MB Formato Adobe PDF Visualizza/Apri	1.23 MB	Adobe PDF	Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11365/1230154