Filtering-Based E-Learning Platforms Extraction using Web Scraping Technique

Web Scraping, E-Learning Platforms, Data Extraction, Online Education, Information Retrieval.

Authors

October 30, 2025

Downloads

Online education, rapidly growing as it is, thus gave rise to a plethora of e-learning platforms such that for any learner, it is becoming increasingly difficult to efficiently identify the best courses. This paper proposes a filtering-based web scraping system that automatically extracts, structures, and presents educational content from various e-learning platforms: Coursera, Udemy, edX, and Khan Academy. Unlike pre-existing methods, our system brings into action real-time scraping with filtration and storage capabilities. This would enable a learner to search through the courses using keywords and categories. The proposed system was engineered using Python (BeautifulSoup, Selenium) and SQL Server with a multithreading engine to boost performance. Experimental results show that the system is capable of extracting more than 5,000 courses in just a few minutes while achieving an average scraping accuracy of 94%. The contributions of this work are threefold: (1) the design of a modular scraping architecture for heterogeneous platforms, (2) implementation of filtering-based aggregation for better course discovery, and (3) evaluation of performance metrics such as response time, scalability, and data accuracy. This research evaluates the potential of automated data aggregation to streamline online education access as well as informed decision-making by the learner.