Deep Learning-Based Vulnerability Detection in Open-Source Software Using Public Datasets
Keywords:
Vulnerability Detection, Deep Learning, Open-Source Software, Public Datasets, Graph Neural Networks, Code Language ModelsAbstract
Automated software vulnerability detection has become a central application of deep learning (DL) in software security, driven by the volume of open-source software (OSS) and the limitations of manual code auditing. This paper reviews the current landscape of DL-based vulnerability detection using public datasets, synthesizing findings from recent surveys, benchmark studies, and dataset papers. It examines the major categories of public vulnerability datasets (synthetic corpora such as SARD/Juliet and real-world, commit-mined corpora such as Devign, Big-Vul, CVEFixes, and DiverseVul), the principal families of DL architectures applied to the task (sequence-based models, graph neural networks, pretrained code language models, and fine-tuned large language models), and the metrics used to evaluate them. The review highlights persistent methodological concerns, including severe class imbalance, label noise inherited from commit-mining pipelines, limited language and vulnerability-type diversity, and a documented generalization gap between in-distribution benchmark performance and realistic, deduplicated, or out-of-distribution evaluation. It concludes that while DL models report strong headline metrics on established benchmarks, current evidence suggests these gains do not always reflect a genuine ability to detect novel vulnerabilities, and it outlines directions for more rigorous dataset construction and benchmarking.
How to cite this article:
Trivedi A K, Dubey S. Deep Learning-Based
Vulnerability Detection in Open-Source Software
Using Public Datasets. J Adv Res Lib Inform Sci
2026; 13(3): 19-23.
DOI: https://doi.org/10.24321/2395.2288.202614
References
Chakraborty, S., Krishna, R., Ding, Y., & Ray, B. (2021). Deep learning-based vulnerability detection: Are we
there yet? IEEE Transactions on Software Engineering, 48(9), 3280–3296.
Chen, Y., Ding, Z., Alowain, L., Chen, X., & Wagner, D. (2023). DiverseVul: A new vulnerable source code dataset for deep learning-based vulnerability detection. In Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses (RAID).
Ding, Y., Fu, Y., Ibrahim, O., Sitawarin, C., Chen, X., Alomair, B., Wagner, D., Ray, B., & Chen, Y. (2024). Vulnerability detection with code language models: How far are we? arXiv preprint arXiv:2403.18624.
Fan, J., Li, Y., Wang, S., & Nguyen, T. N. (2020). A C/ C++ code vulnerability dataset with code changes and CVE summaries (Big-Vul). In Proceedings of the 17th International Conference on Mining Software Repositories.
Harzevili, N. S., Belle, A. B., Wang, J., Wang, S., Jiang, Z. M., & Nagappan, N. (2023). A survey on automated software vulnerability detection using machine learning and deep learning. arXiv preprint arXiv:2306.11673.
National Institute of Standards and Technology. Software Assurance Reference Dataset (SARD). https:// samate.nist.gov/SARD
National Institute of Standards and Technology. National Vulnerability Database (NVD). https://nvd.nist.gov/
Zhou, X., et al. (2024). Multitask-based evaluation of open-source LLM on software vulnerability. arXiv preprint arXiv:2404.02056.
Zhu, Y., et al. (2025). Code vulnerability detection based on augmented program dependency graph and
optimized CodeBERT. Scientific Reports.
Direction for detection: A survey of automated vulnerability detection and all of its pain points. (2024).
arXiv preprint arXiv:2412.11194.
Deep learning aided software vulnerability detection: A survey. (2025). arXiv preprint arXiv:2503.04002.
On benchmarking in machine learning for vulnerability detection. (2025). ISSTA 2025 workshop paper.
Vulnerability datasets for software security: A survey of existing resources, challenges, and future directions.
(2026). ScienceDirect (Journal of Information and Software Technology).