Big Data Analytics Algorithms for Analyzing Employee Leave Patterns in Semarang City using Apache Spark

Apache Spark; Big Data Analytics; Employee Leave; FP-Growth; K-Means Clustering.

Authors

  • Rara Sriartati Redjeki Program Study of Information Systems, Faculty of Information Technology and Industry, Universitas Stikubank, Semarang, Central Java, Indonesia
  • Eko Nur Wahyudi Program Study of Informatics Enginering, Faculty of Information Technology and Industry, Universitas Stikubank, Semarang, Central Java, Indonesia
  • Endang Lestariningsih Program Study of Hospitality, Faculty of Vocational, Universitas Stikubank, Semarang, Central Java, Indonesia
  • Eka Ardhianto Program Study of Information Technology, Faculty of Information Technology and Industry, Universitas Stikubank, Semarang, Central Java, Indonesia
  • Subkhan Indra Gunawan Program Study of Information Technology, Faculty of Information Technology and Industry, Universitas Stikubank, Semarang, Central Java, Indonesia
June 11, 2026
June 13, 2026

Downloads

Leave is an employee right that plays a vital role in maintaining the balance between work productivity and human resource well-being in the workplace. Leave application patterns formed over time can reflect workload dynamics, organizational unit characteristics, and the effectiveness of operational planning. However, in organizations with large numbers of employees and high data volumes, analyzing leave application patterns is often suboptimal when relying on conventional relational database approaches. This study aims to analyze employee leave application patterns in Semarang City by applying a Big Data Analytics approach based on Apache Spark. The dataset used consists of over 150,000 leave application records from the 2020–2025 period, encompassing information on application timing, organizational units, and leave duration. The analysis process was conducted through extraction, transformation, and loading (ETL) stages using Apache Spark DataFrames, including transforming leave data into a daily basis and temporal aggregation. Furthermore, the K-Means algorithm was used to cluster leave application behavior patterns, while FP-Growth was applied to identify frequently occurring combinations of leave timing. The results indicate that most leave applications were made on weekdays with short durations; however, specific clusters showed a tendency to take leave adjacent to or sandwiching public holidays and weekends. These findings demonstrate the existence of recurring and structured leave behavior patterns, which can be utilized as a basis for evaluation and formulation of more effective and adaptive employee leave management policies.