Perumahan dan lingkungan merupakan infrastruktur dasar yang sangat krusial dalam menentukan kualitas hidup dan tingkat kesejahteraan masyarakat. meskipun capaian nasional indikator perumahan di indonesia tergolong tinggi, distribusinya masih menunjukkan heterogenitas yang sangat besar antarwilayah. penelitian ini bertujuan untuk mengelompokkan 508 kabupaten/kota di indonesia berdasarkan tujuh indikator perumahan dan lingkungan menggunakan algoritma clustering large applications (clara) dengan metrik jarak manhattan. jumlah cluster optimal ditentukan melalui kombinasi metode elbow, silhouette, serta empat indeks validitas internal (connectivity, dunn index, davies-bouldin index, dan calinski-harabasz index). hasil penelitian menunjukkan bahwa jumlah kelompok terbaik yang terbentuk adalah 3 cluster. cluster 1 terdiri atas 83 kabupaten/kota (16,34%) yang dikategorikan sebagai wilayah transisional dengan karakteristik ketergantungan yang tinggi terhadap kayu bakar (49,30%) meskipun akses listrik pln sudah memadai (90,40%). cluster 2 merupakan kelompok terbesar yang mendominasi dengan 407 kabupaten/kota (80,12%) dan dikategorikan sebagai wilayah maju karena memiliki profil indikator perumahan dan infrastruktur dasar terbaik. sementara itu, cluster 3 terdiri atas 18 kabupaten/kota (3,54%) yang dikategorikan sebagai wilayah tertinggal (mayoritas berada di wilayah papua) karena memiliki keterbatasan akses listrik pln (13,40%), sanitasi layak (20,00%), serta didominasi oleh rumah tangga dengan luas lantai sempit ≤ 50 m² (84,50%). evaluasi kualitas model menunjukkan nilai dunn index sebesar 0,0923 dan davies-bouldin index sebesar 0,7086, yang menegaskan bahwa algoritma clara mampu menghasilkan struktur pengelompokan yang valid, kompak, dan terpisah secara optimal untuk data berskala besar. hasil tipologi ini diharapkan dapat menjadi landasan pengambilan kebijakan pembangunan infrastruktur yang lebih tepat sasaran berbasis karakteristik wilayah
Electronic Theses and Dissertation
Universitas Syiah Kuala
SKRIPSI
PENGELOMPOKAN KABUPATEN/KOTA DI INDONESIA BERDASARKAN INDIKATOR PERUMAHAN DAN LINGKUNGAN MENGGUNAKAN ALGORITMA CLUSTERING LARGE APPLICATIONS. Banda Aceh Fakultas MIPA Statistika,2026
Baca Juga : ANALISIS FUZZY GEOGRAPHICALLY WEIGHTED CLUSTERING PADA INDIKATOR INDEKS PEMBANGUNAN MANUSIA DI PROVINSI SUMATERA UTARA (Nurul Fadhilah Hayyana Aritonang, 2022)
Abstract
Housing and environment are fundamental infrastructures that are crucial in determining the quality of life and community welfare. Although the national achievements of housing indicators in Indonesia are relatively high, their distribution still shows substantial heterogeneity across regions. This study aims to cluster 508 regencies/cities in Indonesia based on seven housing and environmental indicators using the Clustering Large Applications (CLARA) algorithm with the Manhattan distance metric. The optimal number of clusters was determined through a combination of the Elbow method, Silhouette method, and four internal validity indices (Connectivity, Dunn Index, Davies-Bouldin Index, and Calinski-Harabasz Index). The results indicate that the best cluster structure formed consists of 3 clusters. Cluster 1 comprises 83 regencies/cities (16.34%) categorized as transitional regions, characterized by a high dependency on firewood (49.30%) despite adequate access to PLN electricity (90.40%). Cluster 2 is the largest and dominating group with 407 regencies/cities (80.12%), classified as developed regions due to having the best profile across housing and basic infrastructure indicators. Meanwhile, Cluster 3 consists of 18 regencies/cities (3.54%) categorized as underdeveloped regions (mostly located in the Papua region) due to severely limited access to PLN electricity (13.40%), decent sanitation (20.00%), and a predominance of small floor areas ≤ 50 m² (84.50%). Model evaluation shows a Dunn Index value of 0.0923 and a Davies-Bouldin Index of 0.7086, confirming that the CLARA algorithm successfully produces a valid, compact, and optimally separated cluster structure for large-scale data. This typology is expected to serve as a baseline for data-driven and targeted infrastructure development policies based on regional characteristics.
Baca Juga : VISUALISASI FUZZY CLUSTERING DENGAN MENGGUNAKAN PERANGKAT LUNAK R (Fera Anugreni, 2022)