AN ARTIFICIAL INTELLIGENCE-BASED INFORMATION SYSTEM FOR RESEARCH DUPLICATION DETECTION USING DEEP LEARNING AND DATA CLEANSING TECHNIQUES

Authors

  • Pakorn Kallapadee Faculty of Industrial Technology, Ubon Ratchathani Rajabhat University
  • Piyawat Adthajak Center of Excellence in Artificial Intelligence, Faculty of Industrial Technology, Ubon Ratchathani Rajabhat University
  • Atchariya Laosiri Faculty of Industrial Technology, Ubon Ratchathani Rajabhat University

Keywords:

Information System, Artificial Intelligence, Deep Learning Techniques, Data Cleansing

Abstract

This research focused on the design and development of an information system to detect duplication in research proposal submissions. The system was developed by applying a pre-built model and integrating advanced techniques, specifically Data Cleansing, Text Embedding, and Cosine Similarity. The research objectives were to: 1) develop an information system for detecting research duplication, and 2) evaluate the system's performance. The testing process utilized a dataset of 10 research projects, with the collected data analyzed using semantic similarity criteria, expressed as percentages, and content analysis.

The results indicate that the web-based information system effectively meets user requirements. The implementation of Data Cleansing significantly reduced data noise and enhanced processing accuracy. Furthermore, the application of Text Embedding and Cosine Similarity enabled the system to detect duplication based on contextual and semantic meaning with greater precision than traditional keyword matching methods. Performance evaluations demonstrated that the system could accurately analyze and rank research topic duplication in accordance with established standards. Consequently, the system is suitable for organizational implementation, serving as an effective tool for systematic research data management and supporting administrative decision-making for new research funding allocations, thereby optimizing the management of institutional budgets and research resources.

Downloads

Download data is not yet available.

References

กมลดา เรืองอร่าม. (2566). การพัฒนาและปรับปรุงระบบสารสนเทศเพื่อการจัดการงานวิจัย มหาวิทยาลัยราชภัฏเพชรบุรี. วารสารวิชาการมหาวิทยาลัยราชภัฏเพชรบุรี, 13(3), 12-18.

กรธัช อยู่สุข. (2557). การวิเคราะห์การเผยแพร่ผลงานวิชาการและงานวิจัยซ้ำซ้อนในประเทศไทย (Analysis of Duplicate Publication in Thailand). สืบค้นเมื่อวันที่ 1 เมษายน 2569, จาก https://www.gotoknow.org/posts/579456

บุญชม ศรีสะอาด. (2560). การวิจัยเบื้องต้น (พิมพ์ครั้งที่ 10 ฉบับปรับปรุงใหม่). กรุงเทพฯ: สุวีริยาสาส์น.

ภวิกา อ้วนละมัย, ธรรมนูญ ปัญญาทิพย์, ปนัดดา โพธินาม, อัจฉรา สุมังเกษตร, ณรงค์ฤทธิ์ มะสุใส, ทรงกรด พิมพิศาล, และไพฑูรย์ ทิพย์สันเทียะ. (2568). การพัฒนาระบบแนะนําหนังสือด้วยวิธีการแบบอิงเนื้อหา. วารสารวิศวกรรมและเทคโนโลยีอุตสาหกรรม มหาวิทยาลัยกาฬสินธุ์, 3(3), 36-53.

พีรพัฒน์ หมั่นจิตร์ และฐิตาภรณ์ สินจรูญศักดิ์. (2568). อิทธิพลของคุณภาพระบบสารสนเทศทางการบัญชี คุณภาพงานบริการบัญชีอิเล็กทรอนิกส์ และคุณภาพข้อมูลทางการบัญชีที่ส่งผลต่อประสิทธิผลการปฏิบัติงานด้านบัญชีและความเจริญเติบโตของธุรกิจขนาดกลางและขนาดย่อมในประเทศไทย. วารสารการวิจัยการบริหารการพัฒนา, 15(2), 684-698.

เมทิกา พ่วงแสง และวิสุตา วรรณห้วย. (2562). การพัฒนาระบบสารสนเทศสำหรับการจัดการข้อมูลงานวิจัยในยุคดิจิทัล มหาวิทยาลัยเทคโนโลยีราชมงคลพระนคร. วารสารเทคโนโลยีสื่อสารมวลชน มทร.พระนคร, 4(1), 8-17.

วรพงศ์ บำรุงศรี. (2564). การพัฒนาระบบประมวลผลภาษาธรรมชาติ เพื่อบ่งชี้การกลั่นแกล้งบนโซเชียลมีเดียในภาษาไทย (วิทยานิพนธ์ปริญญามหาบัณฑิต). จุฬาลงกรณ์มหาวิทยาลัย.

สุรพงษ์ วิริยะ, อุทัยวรรณ แก้วตะคุ, Nguyen, H. A., และกิติพิเชษฐ์ ธูปบูชา. (2567). การพัฒนาระบบสารสนเทศเพื่อการสนับสนุนการบริหารจัดการงานวิจัย มหาวิทยาลัยเจ้าพระยา. วารสารวิชาการการประยุกต์ใช้เทคโนโลยีสารสนเทศ, 10(1), 22-35.

เอนก รุ่งนาไร่. (2566). การใช้เทคนิค Data Cleansing เพื่อปรับปรุงคุณภาพข้อมูล. (วิทยานิพนธ์ปริญญามหาบัณฑิต). มหาวิทยาลัยศิลปากร.

เดชชนะ ศรีรัตนลิ้ม และธิดารัตน์ สุขประภาภรณ์. (2568). การพัฒนาระบบสารสนเทศเพื่อการบริหารจัดการงานวิจัยของมหาวิทยาลัยราชภัฏเชียงราย. วารสารการวิจัยกาสะลองคำ มหาวิทยาลัยราชภัฏเชียงราย, 19(1), 58-77.

ธัชพิชญ์ ชำนาญกิจ และฐิติรัตน์ ศิริบวรรัตนกุล. (2565). การตรวจสอบข่าวปลอมภาษาไทยด้วยเทคนิคการประมวลผลภาษาธรรมชาติ. วารสารวิจัยและพัฒนา มจธ., 45(2), 275-287.

โอภาส เอี่ยมสิริวงศ์. (2560). ระบบสารสนเทศเพื่อการจัดการ. กรุงเทพฯ: ซีเอ็ดยูเคชั่น.

Gamma, E., Helm, R., Johnson, R., & Vlissides, J. (1995). Design Patterns: Elements of Reusable Object-Oriented Software. New York: Addison-Wesley.

Manning, C. D., Raghavan, P., & Schütze, H. (2009). An Introduction to Information Retrieval. Cambridge: Cambridge University Press.

Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space. Retrieved from https://arxiv.org/pdf/1301.3781

Sheikh, H., Prins, C., & Schrijvers, E. (2023). Artificial Intelligence: Definition and Background. In Mission AI: Research for Policy. Cham: Springer.

Downloads

Published

2026-06-29

How to Cite

Kallapadee, P., Adthajak, P., & Laosiri, A. (2026). AN ARTIFICIAL INTELLIGENCE-BASED INFORMATION SYSTEM FOR RESEARCH DUPLICATION DETECTION USING DEEP LEARNING AND DATA CLEANSING TECHNIQUES. UMT-Poly Journal, 23(1), 101–116. retrieved from https://so06.tci-thaijo.org/index.php/umt-poly/article/view/295795