PERBANDINGAN METODE SUPPORT VECTOR MACHINE DAN RANDOM FOREST DALAM KLASIFIKASI DAN VISUALISASI TREN PARIWISATA ACEH BERBASIS DATA PLATFORM X | ELECTRONIC THESES AND DISSERTATION

Electronic Theses and Dissertation

Universitas Syiah Kuala

    SKRIPSI

PERBANDINGAN METODE SUPPORT VECTOR MACHINE DAN RANDOM FOREST DALAM KLASIFIKASI DAN VISUALISASI TREN PARIWISATA ACEH BERBASIS DATA PLATFORM X


Pengarang

Zahra Zafira - Personal Name;

Dosen Pembimbing

Kikye Martiwi Sukiakhy - 198605202019032009 - Dosen Pembimbing I
Zulfan - 198606022015041003 - Dosen Pembimbing II



Nomor Pokok Mahasiswa

2208107010040

Fakultas & Prodi

Fakultas MIPA / Informatika (S1) / PDDIKTI : 55201

Subject
-
Kata Kunci
-
Penerbit

Banda Aceh : Fakultas mipa., 2026

Bahasa

No Classification

-

Literature Searching Service

Hard copy atau foto copy dari buku ini dapat diberikan dengan syarat ketentuan berlaku, jika berminat, silahkan hubungi via telegram (Chat Services LSS)

Analisis tren pariwisata secara manual menjadi tidak efisien akibat volume data yang terus meningkat di media sosial seperti Platform X. Penelitian ini bertujuan untuk menerapkan dan membandingkan performa algoritma Support Vector Machine (SVM) dan Random Forest dalam melakukan klasifikasi tren pariwisata Aceh ke dalam empat kategori, yaitu Wisata Alam, Wisata Kuliner, Wisata Budaya dan Rekreasi, dan Non Pariwisata, serta memvisualisasikan hasil klasifikasi menggunakan pustaka Python. Data dikumpulkan menggunakan teknik web scraping dengan tool Tweet-Harvest pada Platform X periode 2024–2025 dan menghasilkan 1.617 tweet awal. Pelabelan data dilakukan secara otomatis berdasarkan kata kunci pencarian (rule-based labeling) yang kemudian divalidasi secara manual, sehingga diperoleh 1.434 data tweet final. Data selanjutnya melalui tahap pra-pemrosesan meliputi case folding, penghapusan URL/mention/hashtag, normalisasi kata tidak baku, penghapusan stopword, tokenisasi, dan stemming menggunakan Sastrawi, kemudian diekstraksi menggunakan metode TF-IDF. Kedua model dilatih menggunakan teknik balanced class weight dengan konfigurasi hyperparameter yang telah ditentukan menggunakan framework Scikit-learn. Hasil evaluasi menunjukkan bahwa model SVM memperoleh performa terbaik dengan nilai Accuracy sebesar 96,52% dan F1-Score sebesar 95,46%, mengungguli Random Forest dengan Accuracy sebesar 96,17% dan F1-Score sebesar 94,92%. Hasil pengujian 5-fold cross validation juga menunjukkan bahwa SVM lebih stabil dengan rata-rata F1-Score sebesar 0,9747 dibandingkan Random Forest sebesar 0,9658. Selain itu, hasil klasifikasi model SVM terhadap 1.434 tweet menunjukkan bahwa Wisata Alam menjadi kategori pariwisata terpopuler (42,9%), diikuti Wisata Kuliner (41,6%), dan Wisata Budaya dan Rekreasi (15,6%).

Kata Kunci: Klasifikasi Teks, Support Vector Machine, Random Forest, TF-IDF, Pariwisata Aceh

Manual analysis of tourism trends has become inefficient due to the increasing volume of data on social media platforms such as Platform X. This study aims to apply and compare the performance of the Support Vector Machine (SVM) and Random Forest algorithms in classifying Aceh tourism trends into four categories: Nature Tourism, Culinary Tourism, Cultural and Recreational Tourism, and Non-Tourism, as well as visualizing the classification results using the Python library. Data were collected using web scraping techniques with the Tweet-Harvest tool on Platform X for the 2024–2025 period, resulting in 1,617 initial tweets. Data labeling was performed automatically based on search keywords (rule-based labeling) and subsequently validated manually, resulting in 1,434 final tweet records. The data were then preprocessed through case folding, removal of URLs/mentions/hashtags, informal word normalization, stopword removal, tokenization, and stemming using Sastrawi, before being extracted using the TF-IDF method. Both models were trained using a balanced class weight technique with predetermined hyperparameter configurations using the Scikit-learn framework. Evaluation results show that the SVM model achieved the best performance with an Accuracy of 96.52% and an F1-Score of 95.46%, outperforming Random Forest, which obtained an Accuracy of 96.17% and an F1-Score of 94.92%. The 5-fold cross validation results also indicate that SVM is more stable, with a mean F1-Score of 0.9747 compared to 0.9658 for Random Forest. Furthermore, the SVM classification results on 1,434 tweets show that Nature Tourism is the most popular tourism category (42.9%), followed by Culinary Tourism (41.6%) and Cultural and Recreational Tourism (15.6%). Keywords: Text Classification, Support Vector Machine, Random Forest, TF-IDF, Aceh Tourism

Citation



    SERVICES DESK