Filtering-Based E-Learning Platforms Extraction using Web Scraping Technique
Downloads
Online education, rapidly growing as it is, thus gave rise to a plethora of e-learning platforms such that for any learner, it is becoming increasingly difficult to efficiently identify the best courses. This paper proposes a filtering-based web scraping system that automatically extracts, structures, and presents educational content from various e-learning platforms: Coursera, Udemy, edX, and Khan Academy. Unlike pre-existing methods, our system brings into action real-time scraping with filtration and storage capabilities. This would enable a learner to search through the courses using keywords and categories. The proposed system was engineered using Python (BeautifulSoup, Selenium) and SQL Server with a multithreading engine to boost performance. Experimental results show that the system is capable of extracting more than 5,000 courses in just a few minutes while achieving an average scraping accuracy of 94%. The contributions of this work are threefold: (1) the design of a modular scraping architecture for heterogeneous platforms, (2) implementation of filtering-based aggregation for better course discovery, and (3) evaluation of performance metrics such as response time, scalability, and data accuracy. This research evaluates the potential of automated data aggregation to streamline online education access as well as informed decision-making by the learner.
S. Patel and M. Gupta, “Automated Data Extraction for Online Education,” IEEE Access, vol. 9, pp. 120340-120352, 2021.
H. Li and J. Wong, “Web Scraping for E-Learning: Opportunities and Challenges,” Computers & Education, vol. 185, 104536, 2022.
Y. Chen, K. Park, and L. Wang, “Course Aggregation Framework for Online Learning,” Information Systems Frontiers, vol. 22, pp. 789-803, 2020.
K. Park and J. Lee, “Data Extraction Challenges in E-Learning Platforms,” Journal of Educational Technology, vol. 17, no. 2, pp. 120-134, 2021.
A. Ahmed and F. Rahman, “Scraping and Recommending MOOCs: A Hybrid Approach,” International Journal of Emerging Technologies in Learning, vol. 17, no. 12, pp. 45-59, 2022.
R. Gupta, T. Singh, and D. Kumar, “Comparative Study of API vs Web Scraping in MOOCs Aggregation,” Procedia Computer Science, vol. 218, pp. 201-209, 2023.
L. Zhou, H. Zhang, and P. Chan, “Ethical Considerations in Educational Web Scraping,” Journal of Information Ethics, vol. 32, no. 1, pp. 35-50, 2023.
S. Kim, Y. Zhou, W. Chen, and Z. He, “NEXT-EVAL: Next evaluation of traditional and LLM web data record extraction,” arXiv preprint arXiv:2505.17125, 2025.
A. Bohra, Y. Zhang, and Y. Yin, “Web Lists: Extracting structured information from complex interactive websites using executable LLM agents,” arXiv preprint arXiv:2504.12682, 2025.
W. Huang, H. Zhang, and X. Sun, “Auto Scraper: A progressive understanding web agent for web scraper generation,” arXiv preprint arXiv:2404.12753, 2024.
