Countering AI-Generated Disinformation: A Novel Detection Model to Safeguard National Security
DOI:
https://doi.org/10.15680/IJCTECE.2025.0804022Keywords:
disinformation detection, AI-generated text, large language models, cross-dataset generalization, shortcut learning, explainable AI, national securityAbstract
Large language models (LLMs) have radically reshaped disinformation from a labor-intensive craft into an easily scalable, fully automated capability, raising very sharp national security, electoral, and trust concerns. Many systems have been proposed, often as automated detectors that achieve nearly perfect within-corpus accuracy; here we report a simple yet effective countermeasure to these victims of automation. In this paper we investigate whether such results lead to directly deployable robustness. We perform cross-corpus deduplication and assemble a single benchmark of 310,065 texts across two similar public corpora; human-written real and fake news (ISOT), a large corpus of news with some GPT-2-generated text, and the DAIGT V2 corpus to evaluate human vs multi-LLM writing (GPT-4, Claude, Llama, Falcon Mistral PaLM). The dual-track pipeline we test contains four classical TF-IDF baselines and a fine-tuned DistilBERT transformer
The macro-F1 and ROC-AUC scores show similar trends: with an overall accuracy of 99.4%, the transformer obtains a macro-F1 of 0.994, while classical baselines get up to macros F1 = 0.937. But under a leave-one-dataset-out (LODO) protocol, performance collapses to near- or below-chance: transformer macro-F1 drops 0.320–0.448, and in the worst-extreme instance ROC-AUC is only 0.327 which is significantly worse than random suggesting learned cues invert across corpora! Explainability analysis (LIME/SHAP and linear-model feature inspection) points to shortcut learning of corpus-specific stylistic and source artifacts agency boilerplate and image-credit tokens, not transferable indicators of deceptive content as the reasons for collapse. We found that within-corpus accuracy is not a reliable indicator of operational performance, and our investigation calls for cross-corpus validation as a prerequisite for deploying disinformation detectors in security-critical settings
References
1. Goldstein, J. A., Sastry, G., Musser, M., DiResta, R., Gentzel, M., & Sedova, K. (2023). Generative Language Models and Automated Influence Operations: Emerging Threats and Potential Mitigations. arXiv:2301.04246.
2. Ahmed, H., Traore, I., & Saad, S. (2017). Detection of Online Fake News Using N-Gram Analysis and Machine Learning Techniques. In Intelligent, Secure, and Dependable Systems in Distributed and Cloud Environments (ISDDC 2017), LNCS 10618, pp. 127–138. Springer.
3. Ahmed, H., Traore, I., & Saad, S. (2018). Detecting Opinion Spams and Fake News Using Text Classification. Security and Privacy, 1(1), e9.
4. Tang, R., Chuang, Y.-N., & Hu, X. (2023). The Science of Detecting LLM-Generated Text. Communications of the ACM (also arXiv:2303.07205).
5. Crothers, E., Japkowicz, N., & Viktor, H. L. (2023). Machine-Generated Text: A Comprehensive Survey of Threat Models and Detection Methods. IEEE Access, 11, 70977–71002.
6. Shu, K., Sliva, A., Wang, S., Tang, J., & Liu, H. (2017). Fake News Detection on Social Media: A Data Mining Perspective. ACM SIGKDD Explorations Newsletter, 19(1), 22–36.
7. Zhou, X., & Zafarani, R. (2020). A Survey of Fake News: Fundamental Theories, Detection Methods, and Opportunities. ACM Computing Surveys, 53(5), 1–40.
8. Wang, W. Y. (2017). “Liar, Liar Pants on Fire”: A New Benchmark Dataset for Fake News Detection. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL), Vol. 2, pp. 422–426.
9. Kaliyar, R. K., Goswami, A., & Narang, P. (2021). FakeBERT: Fake News Detection in Social Media with a BERT-Based Deep Learning Approach. Multimedia Tools and Applications, 80, 11765–11788.
10. Gururangan, S., Swayamdipta, S., Levy, O., Schwartz, R., Bowman, S. R., & Smith, N. A. (2018). Annotation Artifacts in Natural Language Inference Data. In Proceedings of NAACL-HLT, pp. 107–112.
11. Schuster, T., Schuster, R., Shah, D. J., & Barzilay, R. (2020). The Limitations of Stylometry for Detecting Machine-Generated Fake News. Computational Linguistics, 46(2), 499–510.
12. Solaiman, I., Brundage, M., Clark, J., Askell, A., Herbert-Voss, A., Wu, J., Radford, A., et al. (2019). Release Strategies and the Social Impacts of Language Models. arXiv:1908.09203.
13. Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., & Finn, C. (2023). DetectGPT: Zero-Shot Machine-Generated Text Detection Using Probability Curvature. In Proceedings of the 40th International Conference on Machine Learning (ICML), PMLR 202, pp. 24950–24962.
14. Sadasivan, V. S., Kumar, A., Balasubramanian, S., Wang, W., & Feizi, S. (2023). Can AI-Generated Text Be Reliably Detected? arXiv:2303.11156.
15. Pagnoni, A., Graciarena, M., & Tsvetkov, Y. (2022). Threat Scenarios and Best Practices to Detect Neural Fake News. In Proceedings of the 29th International Conference on Computational Linguistics (COLING), pp. 1233–1249.
16. Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., & Wichmann, F. A. (2020). Shortcut Learning in Deep Neural Networks. Nature Machine Intelligence, 2(11), 665–673.
17. Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). “Why Should I Trust You?”: Explaining the Predictions of Any Classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), pp. 1135–1144.
18. Lundberg, S. M., & Lee, S.-I. (2017). A Unified Approach to Interpreting Model Predictions. In Advances in Neural Information Processing Systems (NeurIPS) 30, pp. 4765–4774.
19. Zellers, R., Holtzman, A., Rashkin, H., Bisk, Y., Farhadi, A., Roesner, F., & Choi, Y. (2019). Defending Against Neural Fake News. In Advances in Neural Information Processing Systems (NeurIPS) 32, pp. 9051–9062.
20. Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., & Goldstein, T. (2023). A Watermark for Large Language Models. In Proceedings of the 40th International Conference on Machine Learning (ICML), PMLR 202, pp. 17061–17084.
21. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT, pp. 4171–4186.
22. Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2019). DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter. arXiv:1910.01108.

