Yayın:
Development of Browser Extension for HTML Web Page Content Extraction

dc.contributor.authorKarabulut, Murat
dc.contributor.authorMayda, Islam
dc.date.accessioned2026-06-27T14:32:19Z
dc.date.issued2020
dc.description.abstractAs the amount of content on the websites increases, automatic content extraction from Web pages becomes more important. Although many studies have been done in the literature on this subject, a method that fully solves the problem has not been revealed due to the flexible structure of HTML. The performances of the methods that show success at certain rates also decrease over time with the changing and developing Web structure. In this study, a browser extension was developed to automatically download text content on Web pages. This developed extension provides an output with 100% recall rate by cleaning the text content on the Web page from all tags and codes with a parser that utilizes the Document Object Model (DOM) structure. This browser extension that operates independently from the language has been tested on different types of popular Web sites in Turkey and has been shown to work successfully.en
dc.description.urihttps://doi.org/10.1109/hora49412.2020.9152891
dc.identifier.doi10.1109/hora49412.2020.9152891
dc.identifier.endpage22
dc.identifier.isbn978-1-7281-9352-6
dc.identifier.startpage17
dc.identifier.urihttps://hdl.handle.net/20.500.14981/61802
dc.identifier.wos000644404300002
dc.language.isotur
dc.publisherIEEE
dc.relation.conference2nd International Congress on Human-Computer Interaction, Optimization and Robotic Applications (HORA)
dc.relation.ispartof2ND INTERNATIONAL CONGRESS ON HUMAN-COMPUTER INTERACTION, OPTIMIZATION AND ROBOTIC APPLICATIONS (HORA 2020)
dc.subjectweb content extraction
dc.subjectweb data extraction
dc.subjectweb scraping
dc.subjectComputer Science
dc.subjectRobotics
dc.titleDevelopment of Browser Extension for HTML Web Page Content Extraction
dc.typeProceedings Paper
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar