The Library
Using fuzzy logic to leverage HTML markup for web page representation
Tools
Perez Garcia-Plaza, Alberto, Fresno, Víctor, Martínez, Raquel and Zubiaga, Arkaitz (2017) Using fuzzy logic to leverage HTML markup for web page representation. IEEE Transactions on Fuzzy Systems, 25 (4). pp. 919-933. doi:10.1109/TFUZZ.2016.2586971 ISSN 1063-6706.
|
PDF
WRAP_1373353-cs-280716-tfuzz2586971.pdf - Accepted Version - Requires a PDF viewer. Download (4Mb) | Preview |
Official URL: http://dx.doi.org/10.1109/TFUZZ.2016.2586971
Abstract
The selection of a suitable document representation approach plays a crucial role in the performance of a document clustering task. Being able to pick out representative words within a document can lead to substantial improvements in document clustering. In the case of web documents, the HTML markup that defines the layout of the content provides additional structural information that can be further exploited to identify representative words. In this paper we introduce a fuzzy term weighing approach that makes the most of the HTML structure for document clustering. We set forth and build on the hypothesis that a good representation can take advantage of how humans skim through documents to extract the most representative words. The authors of web pages make use of HTML tags to convey the most important message of a web page through page elements that attract the readers’ attention, such as page titles or emphasized elements. We define a set of criteria to exploit the information provided by these page elements, and introduce a fuzzy combination of these criteria that we evaluate within the context of a web page clustering task. Our proposed approach, called Abstract Fuzzy Combination of Criteria (AFCC), can adapt to datasets whose features are distributed differently, achieving good results compared to other similar fuzzy logic based approaches and TF-IDF across different datasets.
Item Type: | Journal Article | ||||||||
---|---|---|---|---|---|---|---|---|---|
Subjects: | Q Science > QA Mathematics > QA76 Electronic computers. Computer science. Computer software | ||||||||
Divisions: | Faculty of Science, Engineering and Medicine > Science > Computer Science | ||||||||
Library of Congress Subject Headings (LCSH): | Fuzzy systems, Web sites, HTML (Document markup language) | ||||||||
Journal or Publication Title: | IEEE Transactions on Fuzzy Systems | ||||||||
Publisher: | Institute of Electrical and Electronics Engineers | ||||||||
ISSN: | 1063-6706 | ||||||||
Official Date: | August 2017 | ||||||||
Dates: |
|
||||||||
Volume: | 25 | ||||||||
Number: | 4 | ||||||||
Number of Pages: | 17 | ||||||||
Page Range: | pp. 919-933 | ||||||||
DOI: | 10.1109/TFUZZ.2016.2586971 | ||||||||
Status: | Peer Reviewed | ||||||||
Publication Status: | Published | ||||||||
Access rights to Published version: | Restricted or Subscription Access | ||||||||
Date of first compliant deposit: | 29 July 2016 | ||||||||
Date of first compliant Open Access: | 1 August 2016 | ||||||||
Funder: | Spain. Ministerio de Ciencia y Tecnología (MCT), Seventh Framework Programme (European Commission) (FP7) | ||||||||
Grant number: | TIN2013-46616-C2-2-R (MCT), 611233 (FP7) |
Request changes or add full text files to a record
Repository staff actions (login required)
View Item |
Downloads
Downloads per month over past year