Big Data Analytics Algorithms for Analyzing Employee Leave Patterns in Semarang City using Apache Spark
Downloads
Leave is an employee right that plays a vital role in maintaining the balance between work productivity and human resource well-being in the workplace. Leave application patterns formed over time can reflect workload dynamics, organizational unit characteristics, and the effectiveness of operational planning. However, in organizations with large numbers of employees and high data volumes, analyzing leave application patterns is often suboptimal when relying on conventional relational database approaches. This study aims to analyze employee leave application patterns in Semarang City by applying a Big Data Analytics approach based on Apache Spark. The dataset used consists of over 150,000 leave application records from the 2020–2025 period, encompassing information on application timing, organizational units, and leave duration. The analysis process was conducted through extraction, transformation, and loading (ETL) stages using Apache Spark DataFrames, including transforming leave data into a daily basis and temporal aggregation. Furthermore, the K-Means algorithm was used to cluster leave application behavior patterns, while FP-Growth was applied to identify frequently occurring combinations of leave timing. The results indicate that most leave applications were made on weekdays with short durations; however, specific clusters showed a tendency to take leave adjacent to or sandwiching public holidays and weekends. These findings demonstrate the existence of recurring and structured leave behavior patterns, which can be utilized as a basis for evaluation and formulation of more effective and adaptive employee leave management policies.
Michael Armbrust, Reynold S. Xin, Cheng Lian, Yin Huai, Davies Liu, Joseph K. Bradley, Xiangrui Meng, Tomer Kaftan, Michael J. Franklin, Ali Ghodsi, and Matei Zaharia. 2015. Spark SQL: Relational Data Processing in Spark. In Proceedings of the 2015 ACM SIGMOD International Conference on Management of Data (SIGMOD '15). Association for Computing Machinery, New York, NY, USA, 1383–1394.
https://doi.org/10.1145/2723372.2742797
Chen, Hsinchun & Chiang, Roger & Storey, Veda. (2012). Business Intelligence and Analytics: From Big Data to Big Impact. MIS Quarterly. 36. 1165-1188. 10.2307/41703503.
Han, Jiawei & Kamber, Micheline & Pei, Jian. (2012). Data Mining: Concepts and Techniques. 10.1016/C2009-0-61819-5.
OECD. (2020). The OECD Digital Government Policy Framework: Six dimensions of a Digital Government. Paris : OECD Publishing. https://doi.org/10.1787/f64fed2a-en
Hadad, S. H. (2024). Analisis Prioritas Pemberian Cuti Karyawan Menggunakan Metode Pembobotan Entropy dan Simple Multi Attribute Rating Technique. Journal of Artificial Intelligence and Technology Information (JAITI), 2(2), 106-117.
https://doi.org/10.58602/jaiti.v2i2.126
Laksono, et al.. (2025). Perbandingan Apache Airflow dan Apache Spark dalam Proses ETL untuk Memprediksi DropOut dan Keberhasilan Akademik Mahasiswa. Jurnal Informatika dan Rekayasa Perangkat Lunak, 7(2).
Larasati, D. P., Nasrun, M., Si, S., & Ahmad, U. A. (2015). ANALISIS DAN IMPLEMENTASI ALGORITMA FP-GROWTH PADA APLIKASI SMART UNTUK MENENTUKAN MARKET BASKET ANALYSIS PADA USAHA RETAIL ( STUDI KASUS : PT . X ) ANALYSIS AND IMPLEMENTATION OF FP-GROWTH ALGORITHM IN SMART APPLICATION TO DETERMINE MARKET BASKET ANALYSIS ON RETAIL BUSINESS ( CASE STUDY : PT . X ). 2(1), 749–755.
Ines Mergel, Noella Edelmann, Nathalie Haug, Defining digital transformation: Results from expert interviews, Government Information Quarterly, Volume 36, Issue 4, 2019, 101385, ISSN 0740-624X, https://doi.org/10.1016/j.giq.2019.06.002.
Mudjiyanto, Bambang. (2018). TIPE PENELITIAN EKSPLORATIF KOMUNIKASI. Jurnal Studi Komunikasi dan Media. 22. 65.
31445/jskm.2018.220105.
1Elvira Munanda, 2Siti Monalisa, PENERAPAN ALGORITMA FP-GROWTH PADA DATA
TRANSAKSI PENJUALAN UNTUK PENENTUAN TATALETAK, Jurnal Ilmiah Rekayasa dan Manajemen Sistem Informasi, Vol. 7, No. 2, Agustus 2021, Hal. 173-184, e-ISSN 2502-8995 p-ISSN 2460-8181
Azzahra Nur Oktavia, Iqbal Muhammad, Rikky Wahyu Saputra, Muhammad Ichsan Zulfikar, Aries Saifudin. Implementasi Metode Natural Language Processing Dalam Studi Analisis Semantik Dan Emosi Buzzer Pada Tweet Di Aplikasi X, Buletin Ilmiah Ilmu Komputer dan Multimedia (BIIKMA), 2024, pp. 154-159.
Sapta Hastho Ponco*, Susanti, Maziyyatul Mufiedah, A Big Data Approach to Mapping Local Craft Trends: A Study of Creative Economy Opportunities through Online Search Activity in Indonesia, Seminar Nasional Official Statistics 2025.
Juwita, Ita & Ali, Irfan. (2024). PENERAPAN POLA PENJUALAN DENGAN MENGGUNAKAN METODE ALGORITMA ASOSIASI FP-GROWTH BERTUJUAN UNTUK MENINGKATKAN PENJUALAN KOPI DI POINT COFFEE. JATI (Jurnal Mahasiswa Teknik Informatika). 8. 1600-1607. 10.36040/jati.v8i2.9025.
Suryadi, Usep T., and Sri Saraswati. "Sistem Cerdas Pemantau Kenyamanan Ruang Kelas Berbasis Internet of Things (Iot) Menggunakan Metode K-means Pada Platform Thingspeak." Jurnal Teknologi Informasi dan Komunikasi, vol. 15, no. 1, 1 Apr. 2020, pp. 70-81.
Utami, F.D., & Astuti, F.D. (2024). Comparison of Hadoop Mapreduce and Apache Spark in Big Data Processing with Hgrid247-DE. Journal of Applied Informatics and Computing.
Wilda, Roudotul & Saripurna, Darjat & Sulaiman, Oris. (2025). Implementasi Algoritma Frequent Pattern Growth (FP-Growth) untuk Pola Penjualalan Tiket Travel pada PT Taxi Kita Bersama. Hello World Jurnal Ilmu Komputer. 3. 138-145. 10.56211/helloworld.v3i3.588.
Zaharia, Matei & Xin, Reynold & Wendell, Patrick & Das, Tathagata & Armbrust, Michael & Dave, Ankur & Meng, Xiangrui & Rosen, Josh & Venkataraman, Shivaram & Franklin, Michael & Ghodsi, Ali & Gonzalez, Joseph & Shenker, Scott & Stoica, Ion. (2016). Apache spark: A unified engine for big data processing. Communications of the ACM. 59. 56-65. 10.1145/2934664.
