admin@publications.scrs.in   
SCRS Conference Proceedings on Intelligent Systems

Web Scraping Techniques and Applications: A Literature Review

Authors: Chaimaa Lotfi, Swetha Srinivasan, Myriam Ertz and Imen Latrous


Publishing Date: 25-04-2022

ISBN: 978-93-91842-08-6

DOI: https://doi.org/10.52458/978-93-91842-08-6-38

Abstract

Big data analytics gives organizations a way to analyze huge data sets and gather new information. It helps answer basic questions about business operations and business performance. It also helps discover unknown patterns in vast datasets or combinations thereof. In the current data-driven world, it becomes increasingly essential that big data techniques are applied and analyzed for organizational growth. More specifically, with the large availability of data on the Web, whether from social media, websites, online portals, or platforms, to name but a few, it is important for organizations to know how to mine that data in order to extract useful knowledge. Web scraping represents a fundamental approach in this regard. Therefore, this paper aims to provide an updated literature review about the most advanced Web Scraping techniques to better equip scholars and managers with helpful knowledge on how to mine most effectively online data. The paper starts with presenting the basic design of a web scraper and the applications of web scraping in diverse sectors and areas. Next, the different Web scraping methods and Web scraping technologies are presented. Finally, a procedure to develop Web scraping with various tools is proposed before a conclusion wraps up the paper.

Keywords

Big data, web scraping, business performance, web crawling, web mining.

Cite as

Chaimaa Lotfi, Swetha Srinivasan, Myriam Ertz and Imen Latrous, "Web Scraping Techniques and Applications: A Literature Review", In: Raju Pal and Praveen Kumar Shukla (eds), SCRS Conference Proceedings on Intelligent Systems, SCRS, India, 2022, pp. 381-394. https://doi.org/10.52458/978-93-91842-08-6-38

Recent