PENERAPAN ALGORITMA TERM FREQUENCY-INVERSE DOCUMENT FREQUENCY (TF-IDF) UNTUK TEXT MINING
DOI:
https://doi.org/10.30872/jim.v8i3.113Abstract
Algoritma Term Frequency Inverse-Document Frequency merupakan suatu algoritma yang menggalikan antara Term frequency dengan Inverse Document Frequency. Term frequency yaitu jumlah kemunculan sebuah term pada sebuah dokumen. Inverse Document Frequency yaitu pengurangan dominasi term yang sering muncul diberbagai dokumen, dengan memperhitungkan kebalikan frekuensi dokumen yang mengandung suatu kata.
Text Mining pada umumnya adalah unstructured data, atau minimal semistructured. Maka merupakan tantangan tambahan pada text mining yaitu struktur teks yang kompleks dan tidak lengkap, arti yang tidak jelas dan tidak standard, dan bahasa yang berbeda ditambah translasi yang tidak akurat.
Hasil dari penelitian menunjukan bahwa, penerapkan algoritma term frequency inverse-document frequency untuk text mining sangat membantu pengguna. untuk mendapatkan informasi pada kumpulan dokumen. Dengan format file txt berdasarkan kata kunci yang dimasukan oleh pengguna pada sistem. Dengan koleksi uji kata ‘upaya’ pada query maka didapatkan keluaran dengan bobot nilai 8.65441 yang merupakan jumlah kata terbanyak sesuai dengan query.
References
Trunojoyo, H. Sistem Temu Balik Informasi (Sebuah Contoh Implementasi).(http://husni.trunojoyo.ac.id/wpcontent/uploads/2010/03/Husni-IR-dan Klasifikasi.pdf).
Arifin, A. 2002. Penggunaan Digital Tree Hibrida pada Aplikasi Information Retrieval untuk Dokumen Berita. Surabaya : Institut Teknologi Sepuluh Nopember.
Hendry. 2009. Berbagai Aplikasi Databae dengan VB 6.0. Jakarta : PT. Elex Media Komputindo.
Ladjamudin, A. 2005. Analisa dan Desain Sistem Informasi. Yogyakarta : Penerbit Andi Yogyakarta.
Mandala, R. dan Setiawan, H. 2002. Peningkatan Performansi Sistem Temu- Kembali Informasi dengan Perluasan Query Secara Otomatis. Bandung: Institut Teknologi Bandung.
Munawar, 2005. Permodelan Visual dengan UML. Yogyakarta : GRAHA ILMU.
Raymond, J. 2006. Machine Learning Text Categorization. Austin: University of Texas at Austin.
Simarmata, J dan Paryudi, I. 2006. Basis Data. Yogyakarta : Andi.
Naradhipa, R. 2009. Pemilihan Kategori Artikel Berita dengan Text Mining. Paper Terpublikasi. Bandung: Institut Teknologi Bandung.
Ramadhany, T. 2008. Implementasi Kombinasi Model Ruang Vektor dan Model Probabilistik Pada Sistem Temu Balik Informasi. Skripsi Terpublikasi. Bandung: Institut Teknologi Bandung.
http://lecturer.eepis-its.edu/~iwanarif/kuliah/dm/6Text%20Mining.pdf(Tanggal Akses 12 Maret 2011)
http://papers.gunadarma.ac.id/index.php/computer/article/view/574/536(Tanggal Akses 13 Maret 2011)
Downloads
Published
Issue
Section
License
Copyright Transfer StatementThe copyright of this article is transferred to Informatika Mulawarman : Jurnal Ilmiah Ilmu Komputer and when the article is accepted for publication. the authors transfer all and all rights into and to paper including but not limited to all copyrights in the Informatika Mulawarman. The author represents and warrants that the original is the original and that he/she is the author of this paper unless the material is clearly identified as the original source, with notification of the permission of the copyright owner if necessary. The author states that he has the authority and authority to make and carry out this task.
The author states that:
- This paper has not been published in the same form elsewhere.
- This will not be submitted elsewhere for publication prior to acceptance/rejection by this Journal.
A Copyright permission is obtained for material published elsewhere and who require permission for this reproduction. Furthermore, I / We hereby transfer the unlimited publication rights of the above paper to Informatika Mulawarman : Jurnal Ilmiah Ilmu Komputer. Copyright transfer includes exclusive rights to reproduce and distribute articles, including reprints, translations, photographic reproductions, microforms, electronic forms (offline, online), or other similar reproductions.
The author's mark is appropriate for and accepts responsibility for releasing this material on behalf of any and all coauthor. This Agreement shall be signed by at least one author who has obtained the consent of the co-author (s) if applicable. After the submission of this agreement is signed by the author concerned, the amendment of the author or in the order of the author listed shall not be accepted.
Rights / Terms and Conditions Saved
- The author keeps all proprietary rights in every process, procedure, or article creation described in Work.
- The author may reproduce or permit others to reproduce the work or derivative works for the author's personal use or for the use of the company, provided that the source and the Informatika Mulawarman copyright notice are indicated, the copy is not used in any way implying the Journal of Informatika Mulawarman (JIM) approval of the product or service from any company, and the copy itself is not offered for sale.
- Although authors are permitted to reuse all or part of the Works in other works, this does not include granting third-party requests to reprint, republish, or other types of reuse.

Informatika Mulawarman by http://e-journals.unmul.ac.id/index.php/JIM/index is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
under the CC BY-SA license, authors and other users are able to reprint, distribute or use the material for commercial purposes so long as they give attribution to the journal Informatika Mulawarman and license the republished material under the same license.