Detecting AI-Generated Scientific Abstracts Using Galactica and Graph Neural Networks
Date
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
The rise of large language models has introduced new challenges in maintaining research integrity, particularly through the potential proliferation of AI-generated scientific content. This paper presents a novel hybrid framework that combines Galactica-a scientific language model developed by Meta AI-with Graph Neural Networks (GNNs) to detect AI-generated research abstracts. Leveraging the AI-GA dataset comprising 28,662 labeled abstracts, we extract domain-specific semantic embeddings using Galactica and construct a semantic similarity graph to approximate citation-like relationships. A two- layer GCN is then trained to classify each abstract as human- or AI-authored. Experimental results demonstrate that our method outperforms traditional baselines such as TF-IDF, RoBERTa, and perplexity-based detectors, while offering interpretable and scalable detection suitable for editorial screening pipelines.
