Detecting near-replicas on the Web by content and hyperlink analysis

IRIS

The presence of near-replicas of documents is very common on the Web. Documents may be replicated completely or partially for different reasons (versions, mirrors, etc.), or the same resource can be associated to different URLs (dynamically generated pages, etc.). Whilst replication can improve information accessibility by the users, the presence of near-replicated documents can hinder the effectiveness of search engines (for example, decreasing the coverage). We propose a method to detect similar pages, in particular replicas and near-replicas, which is based on a pair of signatures. The first signature is obtained by a random projection of the bag-of-words vector representing the page contents. The second signature is computed by a recursive equation which exploits the connectivity among the Web pages to code the context of each page. The accuracy of the proposed approach is analyzed and validated by experimental results which show that on the given dataset near-replicas can be detected with a precision-recall of 93%.

Di Iorio, E., Diligenti, M., Gori, M., Maggini, M., Pucci, A. (2003). Detecting near-replicas on the Web by content and hyperlink analysis. In Proceedings of the IEEE/WIC International Conference on Web Intelligence (WIC 2003) (pp.249-255). Los Alamitos, USA : IEEE COMPUTER SOC [10.1109/WI.2003.1241201].

Detecting near-replicas on the Web by content and hyperlink analysis

Di Iorio E.;Diligenti M.;Gori M.;Maggini M.;Pucci A.

2003-01-01

Abstract

The presence of near-replicas of documents is very common on the Web. Documents may be replicated completely or partially for different reasons (versions, mirrors, etc.), or the same resource can be associated to different URLs (dynamically generated pages, etc.). Whilst replication can improve information accessibility by the users, the presence of near-replicated documents can hinder the effectiveness of search engines (for example, decreasing the coverage). We propose a method to detect similar pages, in particular replicas and near-replicas, which is based on a pair of signatures. The first signature is obtained by a random projection of the bag-of-words vector representing the page contents. The second signature is computed by a recursive equation which exploits the connectivity among the Web pages to code the context of each page. The accuracy of the proposed approach is analyzed and validated by experimental results which show that on the given dataset near-replicas can be detected with a precision-recall of 93%.

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno
	
				2003
			
	Codice ISBN
	
				0769519326
			
	Citazione
	
				Di Iorio, E., Diligenti, M., Gori, M., Maggini, M., Pucci, A. (2003). Detecting near-replicas on the Web by content and hyperlink analysis. In Proceedings of the IEEE/WIC International Conference on Web Intelligence (WIC 2003) (pp.249-255). Los Alamitos, USA : IEEE COMPUTER SOC [10.1109/WI.2003.1241201].
			
	Appare nelle tipologie:
	
				4.1 Contributo in Atti di convegno

File in questo prodotto:

File	Dimensione	Formato
WI03.pdf non disponiibile Tipologia: PDF editoriale Licenza: NON PUBBLICO - Accesso privato/ristretto Dimensione 265.56 kB Formato Adobe PDF Visualizza/Apri Richiedi una copia	265.56 kB	Adobe PDF	Visualizza/Apri Richiedi una copia

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11365/36822

Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo