Kanker paru-paru adenokarsinoma sering terdeteksi pada tahap lanjut sehingga analisis biomarker dna diperlukan untuk membantu mengenali perubahan genetik yang berkaitan dengan kanker. penelitian ini bertujuan membangun model klasifikasi sekuens dna, membandingkan kinerja bidirectional long short term memory (bilstm) dan bidirectional gated recurrent unit (bigru), serta menentukan konfigurasi yang lebih baik dalam membedakan sekuens normal dan kanker. data sekunder diperoleh dari ncbi dan terdiri atas 273 sekuens dari gen egfr, alk, braf, pd-l1, dan ros1. data diproses melalui penyaringan, penyejajaran sekuens, pembagian berdasarkan sumber berkas dengan rasio 70% data latih, 10% data validasi, dan 20% data uji, augmentasi substitusi kodon sinonim, tokenisasi, serta penyamaan panjang sekuens. dua skenario augmentasi digunakan, yaitu target 500 dan 100 sekuens per gen. model dilatih menggunakan binary focal crossentropy, bobot kelas, dan early stopping. evaluasi dilakukan menggunakan akurasi, presisi, recall, f1-score, confusion matrix, dan validasi pada sekuens eksternal nr_103548.1. hasil pengujian menunjukkan bahwa bigru pada skenario augmentasi 500 memberikan kinerja terbaik dengan test loss 0,2671, akurasi 92,45%, presisi kanker 100%, recall kanker 84,62%, serta f1-score 0,92 untuk kelas kanker dan 0,93 untuk kelas normal. model tersebut menghasilkan empat false negative, lebih sedikit daripada bilstm yang menghasilkan sembilan false negative dengan akurasi 81,13%. pada skenario augmentasi 100, kedua model memperoleh akurasi 64,10% dan gagal mengenali seluruh 13 sampel normal, sehingga menunjukkan bias yang kuat terhadap kelas kanker. pada validasi eksternal skenario augmentasi 500, bigru memprediksi kanker dengan tingkat keyakinan 78,96%, lebih tinggi daripada bilstm sebesar 59,32%. berdasarkan hasil tersebut, bigru dengan skenario augmentasi 500 merupakan konfigurasi terbaik dalam penelitian ini. namun, model belum dapat digunakan sebagai alat pemeriksaan medis karena masih menghasilkan false negative dan baru diuji pada satu sekuens eksternal. kata kunci: adenokarsinoma paru-paru, biomarker dna, klasifikasi sekuens, bilstm, bigru, augmentasi data.
Electronic Theses and Dissertation
Universitas Syiah Kuala
SKRIPSI
DETEKSI KANKER PARU PARURNADENOKARSINOMA DENGAN ALGORITMARNRECURRENT NEURAL NETWORK (RNN)RNMENGGUNAKAN DATA BIOMARKER DNA. Banda Aceh Fakultas mipa,2026
Baca Juga : DETEKSI DAN MENGHITUNG LUAR AREA KANKER PADA CITRA CT-SCAN ORGAN PARU MENGGUNAKAN MATLAB (MAYA ANDINA, 2020)
Abstract
Lung adenocarcinoma is frequently diagnosed at an advanced stage, highlighting the need for DNA biomarker analysis to support the identification of cancer-related genetic alterations. This study developed a DNA sequence classification model, compared the performance of Bidirectional Long Short Term Memory (BiLSTM) and Bidirectional Gated Recurrent Unit (BiGRU), and identified the more effective configuration for distinguishing normal and cancer-associated sequences. A secondary dataset containing 273 sequences from the EGFR, ALK, BRAF, PD-L1, and ROS1 genes was obtained from NCBI. The data preparation process included sequence filtering and alignment, a source-grouped split into 70% training, 10% validation, and 20% testing data, synonymous codon substitution, tokenization, and sequence padding. Two augmentation settings were evaluated, targeting 500 and 100 sequences per gene. The models were trained using Binary Focal Crossentropy, class weighting, and early stopping. Performance was assessed using accuracy, precision, recall, F1-score, a confusion matrix, and external validation on sequence NR_103548.1. The results showed that BiGRU under the 500-sequence augmentation setting achieved the best overall performance, with a test loss of 0.2671, an accuracy of 92.45%, a cancer-class precision of 100%, a cancer-class recall of 84.62%, and F1-scores of 0.92 and 0.93 for the cancer and normal classes, respectively. BiGRU produced four False Negatives, compared with nine from BiLSTM, which achieved an accuracy of 81.13%. Under the 100-sequence setting, both models achieved 64.10% accuracy but failed to identify all 13 normal samples, indicating a strong prediction bias toward the cancer class. In external validation using the 500-sequence setting, BiGRU predicted cancer with 78.96% confidence, compared with 59.32% for BiLSTM. These findings identify BiGRU with the 500-sequence augmentation setting as the best configuration evaluated in this study. However, the model is not yet suitable for medical screening because it still produces False Negatives and has been externally evaluated on only one sequence. Keywords: lung adenocarcinoma, DNA biomarkers, sequence classification, BiLSTM, BiGRU, data augmentation.
Baca Juga : DETEKSI DINI KANKER PAYUDARA BERBASIS TERMOGRAFI DAN DEEP LEARNING (Roslidar, 2022)