LLM-Guided Semantic Rule Mining for Automated Cybersecurity Threat Intelligence Extraction

Authors

  • Cody Gustafsson Department of Computer Science, University of Houston, Houston, TX, USA. Author
  • Anish K. Nukharjee Department of Computer Science, University of New Hampshire, Durham, NH, USA. Author
  • Cady Nershell Department of Electrical Engineering and Computer Science, University of Kansas, Lawrence, KS, USA. Author
  • Hugo White Department of Computer Science, University of Central Florida, Orlando, FL, USA. Author

Keywords:

cybersecurity threat intelligence, large language models, semantic rule mining, knowledge extraction, automated reasoning, system architecture

Abstract

The accelerating sophistication of cyber threats and the exponential growth of unstructured threat data demand a paradigm shift in how actionable intelligence is extracted, synthesized, and operationalized. Traditional rule-based systems and manual curation fail to scale, while purely data-driven approaches often produce brittle patterns lacking semantic depth or contextual awareness. This paper presents a comprehensive system-level investigation into LLM-guided semantic rule mining, a novel architectural paradigm that couples the generative reasoning capabilities of large language models with structured rule discovery mechanisms to automate the extraction of cybersecurity threat intelligence. We examine the end-to-end infrastructure, from ingestion of heterogeneous threat feeds, through semantic enrichment and embedding, to the iterative refinement of human-interpretable rules that capture adversary behaviors, indicators of compromise, and attack patterns. The discussion foregrounds structural trade-offs among latency, accuracy, explainability, and adversarial robustness, and analyzes how modular pipeline design can balance real-time demands with forensic depth. Governance considerations, including bias propagation, fairness across threat actor attribution, and the sustainability of large-scale language model deployments, are treated as first-order architectural constraints rather than afterthoughts. Through detailed conceptual analysis, cross-domain comparisons, and forward-looking perspectives on policy alignment, we establish a blueprint for deploying LLM-guided semantic rule mining as a foundational component of next-generation security operations centers and national cyber defense infrastructures.

References

1. Strom, B. E., Applebaum, A., Miller, D. P., Nickels, K. C., Pennington, A. G., & Thomas, C. B. (2018). MITRE ATT&CK: Design and philosophy. MITRE Corporation.

2. Barnum, S. (2014). Standardizing cyber threat intelligence information with the Structured Threat Information Expression (STIX). MITRE Corporation.

3. Aggarwal, C. C., & Han, J. (Eds.). (2014). Frequent pattern mining. Springer.

4. Galárraga, L. A., Teflioudi, C., Hose, K., & Suchanek, F. M. (2013). AMIE: Association rule mining under incomplete evidence in ontological knowledge bases. Proceedings of the 22nd International Conference on World Wide Web, 413-422.

5. Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., ... & Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877-1901.

6. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.

7. Bayer, M., Kuehn, P., Shanehsazzadeh, A., & Molloy, I. (2022). SecureBERT: A domain-specific language model for cybersecurity. Proceedings of the 15th ACM Conference on Security and Privacy in Wireless and Mobile Networks, 39-49.

8. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., ... & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459-9474.

9. Al-Shaer, E., Duan, Q., & Jajodia, S. (2015). Automated security intelligence generation using network threat graphs. In Network security (pp. 287-306). Springer.

10. Mittal, S., Das, A. K., Mulwad, V., Joshi, A., & Finin, T. (2016). CyberTwitter: Using Twitter to generate alerts for cybersecurity threats and vulnerabilities. 2016 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM), 860-867.

11. Wang, Q., Mao, Z., Wang, B., & Guo, L. (2017). Knowledge graph embedding: A survey of approaches and applications. IEEE Transactions on Knowledge and Data Engineering, 29(12), 2724-2743.

12. Liao, X., Yuan, K., Wang, X., Li, Z., Xing, L., & Beyah, R. (2016). Acing the IOC game: Toward automatic discovery and analysis of open-source cyber threat intelligence. Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, 755-766.

13. Shu, X., Tian, K., & Yao, D. D. (2018). Breaking the target: An in-depth analysis of threat actor attribution and cyber campaign characterization. IEEE Security & Privacy, 16(5), 44-53.

14. Han, Z., Chen, W., Han, Y., Mao, R., & Qin, J. (2026). Fast Diversified Top-k Rule Discovery via User-Guided Embeddings. IEEE Transactions on Knowledge and Data Engineering.

15. Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 30.

16. Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 1-35.

17. Bonawitz, K., Ivanov, V., Kreuter, B., Marcedone, A., McMahan, H. B., Patel, S., ... & Seth, K. (2017). Practical secure aggregation for privacy-preserving machine learning. Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, 1175-1191.

18. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 3645-3650.

19. National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework (NIST AI 100-1). U.S. Department of Commerce.

20. European Commission. (2021). Proposal for a Regulation of the European Parliament and of the Council laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). COM/2021/206 final.

Downloads

Published

2026-06-09

How to Cite

LLM-Guided Semantic Rule Mining for Automated Cybersecurity Threat Intelligence Extraction. (2026). Journal of Data Intelligence and AI Systems, 1(3). https://www.jdataai.org/index.php/home/article/view/110