Web informative content identification and filtering using machine learning technique. (2016)
- Record Type:
- Journal Article
- Title:
- Web informative content identification and filtering using machine learning technique. (2016)
- Main Title:
- Web informative content identification and filtering using machine learning technique
- Authors:
- Narwal, Neetu
Sharma, Sanjay Kumar - Abstract:
- Internet has gained greatest acceptance as reservoirs of information. It has been observed that the web page along with main content comprises of noise (advertisement, external links), which poses difficulty for various search engines crawlers to correctly classify the web page and it also provides distraction to the user interested in gathering relevant data. In this paper, we proposed a novel approach which categorises the relevant content from the web page and use this information to filter and rearrange the content of the web page. We used the web page segmentation algorithm for parsing the web page to obtain non-overlapping visual blocks and then extracted the features from these visual blocks to build the dataset. The dataset have been trained using popular machine learning classifier techniques (neural network, RBF neural network) to discriminate content. Finally, the classification output is used to perform main content filtering of the web page. We also analysed the importance of features on the learning process and perceive that the embedded objects from external source have highest significance for block identification.
- Is Part Of:
- International journal of data analysis techniques and strategies. Volume 8:Number 4(2016)
- Journal:
- International journal of data analysis techniques and strategies
- Issue:
- Volume 8:Number 4(2016)
- Issue Display:
- Volume 8, Issue 4 (2016)
- Year:
- 2016
- Volume:
- 8
- Issue:
- 4
- Issue Sort Value:
- 2016-0008-0004-0000
- Page Start:
- 332
- Page End:
- 347
- Publication Date:
- 2016
- Subjects:
- web information retrieval -- page segmentation -- visual blocks -- embedded objects -- web content identification -- web content filtering -- machine learning -- webpage contents -- feature extraction -- neural networks -- block identification
Electronic data processing -- Periodicals
Database searching -- Periodicals
005 - Journal URLs:
- http://www.inderscience.com/jhome.php?jcode=ijdats ↗
http://www.inderscience.com/ ↗ - Languages:
- English
- ISSNs:
- 1755-8050
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library STI - ELD Digital store - Ingest File:
- 8143.xml