ALGORITMA PREDIKSI OUTLIER MENGGUNAKAN BORDER SOLVING SET
DOI:
https://doi.org/10.30872/jim.v9i3.172Keywords:
AnalisisOutlier, PrediksiOutlier, Data Tepi Klaster, Solving Set, DataMiningAbstract
Prediksi outlier penting untuk menjaga validitas data. Algoritma prediksi outlier konvensional memiliki kelemahan dalam hal efisiensi karena harus membandingkan data yang akan diprediksi dengan seluruh data dalam data set. Konsep baru yang melibatkan solving setmuncul sebagai solusi atas permasalahan efisiensi dalam prediksi outlier. Dengan menggunakan solving set, waktu prediksi menjadi lebih cepat tetapi akurasi
prediksi menjadi lebih jelek. Dalam penelitian ini dikembangkan suatu algoritma prediksi outlier baru yang
efisien dalam melakukan prediksi tetapi tidak mengorbankan akurasi hasil prediksi. Algoritma baru ini
merupakan inovasi terhadap konsep solving setyang sudah dikembangkan sebelumnya. Dalam penelitian
sebelumnya, solving setdidefinisikan sebagai subset dari data set yang beranggotakan data yang menjadi top n-outlier sebagai representasi data set. Sedangkan dalam penelitian ini, solving setdidefinisikan ulang sebagai
subset dari data set yang merupakan data tepi klaster beserta pusat klasternya sebagai representasi data set, atau selanjutnya disebut border solving set. Data tepi klaster dideteksi menggunakan algoritmaBORDER yang telah terbukti dapat mendeteksi data tepi klaster secara efisien, dan algoritma klasterisasi berbasis hirarki digunakan untuk melakukan klasterisasi data tepi yang telah terdeteksi. Selanjutnya, pusat masing-masing klaster dicari dengan menghitung nilai median dari data tepi pada masing-masing klaster. Algoritma Prediksi Outlier dalam penelitian ini dilakukan dengan membandingkan jarakantara data yang akan diprediksi (query data) dengan pusat klaster dan jarak antara query datadengan data tepi klaster yang terdekat. Algoritma Prediksi Outlier pada penelitian ini selanjutnya disebut APOTEK (Algoritma Prediksi Outlier menggunakan TEpi Klaster). Setelah dilakukan beberapa percobaan terhadap beberapa dataset dengan distribusi normal dan seragam, APOTEK terbukti dapat melakukan perbaikan terhadap algoritma prediksi outlier yang sudah ada sebelumnya. Dalam aspek akurasi prediksi, APOTEK berhasil melakukan peningkatan sebesar 5% dibandingkan dengan algoritma prediksi outlier yang dikembangkan oleh Angiulli et. al. (2006), untuk data set berdistribusi normal.
References
Chenyi Xia, Wynne Hsu, Mong Li Lee dan Beng Chin Ooi, (2006), "BORDER: Efficient Computation of Boundary Points”, IEEE Transactions on Knowledge and Data Engineering, vol. 18, no. 3, hal. 289-303.
Fabrizio Angiulli, Stefano Basta dan Clara Pizzutti, (2006), “Distance-Based Detection and Prediction of Outliers”, IEEE Transactions on Knowledge and Data Engineering, vol. 18, no. 2, hal. 145-160.
Hui Xiong, Gaurav Pandey, Michael Steinbach dan Vipin Kumar, (2006), “Enhancing Data Analysis with Noise Removal”, IEEE Transactions on Knowledge and Data Engineering, vol. 18, no. 3, hal. 304-319.
Lubsa, Dana Avram, (2005), “Unsupervised Single-Link Hierarchical Clustering”, Studia Univ. Babes-Bolyai, Informatica, vol. 1, no. 2.
Pang-Ning Tan, Michael Steinbach dan Vipin Kumar, (2006), Introduction to Data Mining, Pearson Education, Inc., Boston.
Downloads
Published
Issue
Section
License
Copyright Transfer StatementThe copyright of this article is transferred to Informatika Mulawarman : Jurnal Ilmiah Ilmu Komputer and when the article is accepted for publication. the authors transfer all and all rights into and to paper including but not limited to all copyrights in the Informatika Mulawarman. The author represents and warrants that the original is the original and that he/she is the author of this paper unless the material is clearly identified as the original source, with notification of the permission of the copyright owner if necessary. The author states that he has the authority and authority to make and carry out this task.
The author states that:
- This paper has not been published in the same form elsewhere.
- This will not be submitted elsewhere for publication prior to acceptance/rejection by this Journal.
A Copyright permission is obtained for material published elsewhere and who require permission for this reproduction. Furthermore, I / We hereby transfer the unlimited publication rights of the above paper to Informatika Mulawarman : Jurnal Ilmiah Ilmu Komputer. Copyright transfer includes exclusive rights to reproduce and distribute articles, including reprints, translations, photographic reproductions, microforms, electronic forms (offline, online), or other similar reproductions.
The author's mark is appropriate for and accepts responsibility for releasing this material on behalf of any and all coauthor. This Agreement shall be signed by at least one author who has obtained the consent of the co-author (s) if applicable. After the submission of this agreement is signed by the author concerned, the amendment of the author or in the order of the author listed shall not be accepted.
Rights / Terms and Conditions Saved
- The author keeps all proprietary rights in every process, procedure, or article creation described in Work.
- The author may reproduce or permit others to reproduce the work or derivative works for the author's personal use or for the use of the company, provided that the source and the Informatika Mulawarman copyright notice are indicated, the copy is not used in any way implying the Journal of Informatika Mulawarman (JIM) approval of the product or service from any company, and the copy itself is not offered for sale.
- Although authors are permitted to reuse all or part of the Works in other works, this does not include granting third-party requests to reprint, republish, or other types of reuse.

Informatika Mulawarman by http://e-journals.unmul.ac.id/index.php/JIM/index is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
under the CC BY-SA license, authors and other users are able to reprint, distribute or use the material for commercial purposes so long as they give attribution to the journal Informatika Mulawarman and license the republished material under the same license.