Skip to content Skip to navigation
University of Warwick
  • Study
  • |
  • Research
  • |
  • Business
  • |
  • Alumni
  • |
  • News
  • |
  • About

University of Warwick
Publications service & WRAP

Highlight your research

  • WRAP
    • Home
    • Search WRAP
    • Browse by Warwick Author
    • Browse WRAP by Year
    • Browse WRAP by Subject
    • Browse WRAP by Department
    • Browse WRAP by Funder
    • Browse Theses by Department
  • Publications Service
    • Home
    • Search Publications Service
    • Browse by Warwick Author
    • Browse Publications service by Year
    • Browse Publications service by Subject
    • Browse Publications service by Department
    • Browse Publications service by Funder
  • Help & Advice
University of Warwick

The Library

  • Login
  • Admin

Using fuzzy logic to leverage HTML markup for web page representation

Tools
- Tools
+ Tools

Perez Garcia-Plaza, Alberto, Fresno, Víctor, Martínez, Raquel and Zubiaga, Arkaitz (2017) Using fuzzy logic to leverage HTML markup for web page representation. IEEE Transactions on Fuzzy Systems, 25 (4). pp. 919-933. doi:10.1109/TFUZZ.2016.2586971

[img]
Preview
PDF
WRAP_1373353-cs-280716-tfuzz2586971.pdf - Accepted Version - Requires a PDF viewer.

Download (4Mb) | Preview
Official URL: http://dx.doi.org/10.1109/TFUZZ.2016.2586971

Request Changes to record.

Abstract

The selection of a suitable document representation approach plays a crucial role in the performance of a document clustering task. Being able to pick out representative words within a document can lead to substantial improvements in document clustering. In the case of web documents, the HTML markup that defines the layout of the content provides additional structural information that can be further exploited to identify representative words. In this paper we introduce a fuzzy term weighing approach that makes the most of the HTML structure for document clustering. We set forth and build on the hypothesis that a good representation can take advantage of how humans skim through documents to extract the most representative words. The authors of web pages make use of HTML tags to convey the most important message of a web page through page elements that attract the readers’ attention, such as page titles or emphasized elements. We define a set of criteria to exploit the information provided by these page elements, and introduce a fuzzy combination of these criteria that we evaluate within the context of a web page clustering task. Our proposed approach, called Abstract Fuzzy Combination of Criteria (AFCC), can adapt to datasets whose features are distributed differently, achieving good results compared to other similar fuzzy logic based approaches and TF-IDF across different datasets.

Item Type: Journal Article
Subjects: Q Science > QA Mathematics > QA76 Electronic computers. Computer science. Computer software
Divisions: Faculty of Science, Engineering and Medicine > Science > Computer Science
Library of Congress Subject Headings (LCSH): Fuzzy systems, Web sites, HTML (Document markup language)
Journal or Publication Title: IEEE Transactions on Fuzzy Systems
Publisher: Institute of Electrical and Electronics Engineers
ISSN: 1063-6706
Official Date: August 2017
Dates:
DateEvent
August 2017Published
7 July 2016Available
6 June 2016Accepted
Volume: 25
Number: 4
Number of Pages: 17
Page Range: pp. 919-933
DOI: 10.1109/TFUZZ.2016.2586971
Status: Peer Reviewed
Publication Status: Published
Access rights to Published version: Restricted or Subscription Access
Funder: Spain. Ministerio de Ciencia y Tecnología (MCT), Seventh Framework Programme (European Commission) (FP7)
Grant number: TIN2013-46616-C2-2-R (MCT), 611233 (FP7)

Request changes or add full text files to a record

Repository staff actions (login required)

View Item View Item

Downloads

Downloads per month over past year

View more statistics

twitter

Email us: wrap@warwick.ac.uk
Contact Details
About Us