Abstract
Web Page Noise Cleaning is one of the new research areas of study for removing the noise patterns of web pages for effective web mining. The World Wide Web contains large amount of web pages which are accessible to users. With conventional data or text, Web pages generally contain a large amount of noise information that is not part of the main contents of the web pages, e.g., advertisement banners, navigation bars, and disclaimer/copyright notices. The main objective of this area is removing such irrelevant information (i.e. Web Page Noise or Local Noise) in Web pages that can seriously harm Web mining task such as clustering and classification etc. For detection and removal of noises a new DOM tree structure is proposed. After DOM tree construction, we can implement DUSTER framework for crawling the document using normalized rules. The result shows the remarkable increase in F score and accuracy is obtained. In this work, we focus on detecting and eliminating local noises in Web pages to improve the performance of Web mining that is Web page clustering and classification. Then our experimental results show that improved performance at the time of classification and clustering.
Keywords
Noise cleaning
DOM tree
DUSTER
web mining
clustering
classification
Authors
How to Cite this Article
E. Silambarasan, M.Sindhuja, T.Pragathi, S.Siva Abbirammi, K.Gayathri (2016).
"REMOVAL OF MALICIOUS SIDE INFORMATION IN WEB DOCUMENTS".
International Journal of Contemporary Research in Computer Science and Technology,
2(3), pp. 551-554.