A Short Survey on Deep Learning for Multimodal Integration: Applications, Future Perspectives and Challenges

IRIS

Deep learning has achieved state-of-the-art performances in several research applications nowadays: from computer vision to bioinformatics, from object detection to image generation. In the context of such newly developed deep-learning approaches, we can define the concept of multimodality. The objective of this research field is to implement methodologies which can use several modalities as input features to perform predictions. In this, there is a strong analogy with respect to what happens with human cognition, since we rely on several different senses to make decisions. In this article, we present a short survey on multimodal integration using deep-learning methods. In a first instance, we comprehensively review the concept of multimodality, describing it from a two-dimensional perspective. First, we provide, in fact, a taxonomical description of the multimodality concept. Secondly, we define the second multimodality dimension as the one describing the fusion approaches in multimodal deep learning. Eventually, we describe four applications of multimodal deep learning to the following fields of research: speech recognition, sentiment analysis, forensic applications and image processing.

Dimitri, G.M. (2022). A Short Survey on Deep Learning for Multimodal Integration: Applications, Future Perspectives and Challenges. COMPUTERS, 11(11) [10.3390/computers11110163].

A Short Survey on Deep Learning for Multimodal Integration: Applications, Future Perspectives and Challenges

Dimitri, Giovanna Maria

2022-01-01

Abstract

Deep learning has achieved state-of-the-art performances in several research applications nowadays: from computer vision to bioinformatics, from object detection to image generation. In the context of such newly developed deep-learning approaches, we can define the concept of multimodality. The objective of this research field is to implement methodologies which can use several modalities as input features to perform predictions. In this, there is a strong analogy with respect to what happens with human cognition, since we rely on several different senses to make decisions. In this article, we present a short survey on multimodal integration using deep-learning methods. In a first instance, we comprehensively review the concept of multimodality, describing it from a two-dimensional perspective. First, we provide, in fact, a taxonomical description of the multimodality concept. Secondly, we define the second multimodality dimension as the one describing the fusion approaches in multimodal deep learning. Eventually, we describe four applications of multimodal deep learning to the following fields of research: speech recognition, sentiment analysis, forensic applications and image processing.

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno
	
				2022
			
	Rivista su cui è pubblicata l'opera
	
				COMPUTERS
			
	Citazione
	
				Dimitri, G.M. (2022). A Short Survey on Deep Learning for Multimodal Integration: Applications, Future Perspectives and Challenges. COMPUTERS, 11(11) [10.3390/computers11110163].
			
	Appare nelle tipologie:
	
				1.1 Articolo in rivista

File in questo prodotto:

File	Dimensione	Formato
computers-11-00163.pdf accesso aperto Descrizione: Paper Tipologia: PDF editoriale Licenza: Creative commons Dimensione 575.09 kB Formato Adobe PDF Visualizza/Apri	575.09 kB	Adobe PDF	Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11365/1220301