AI-Enhanced Biomedical Literature Retrieval via Large Language Models and Researcher Preference-Aware Semantic Matching

Authors

  • Anil Mehra Department of Computer Science and Engineering, University of Nevada, Reno, Reno, NV, USA. Author
  • Martins C. Russell Department of Computer Science, George Mason University, Fairfax, VA, USA. Author

Keywords:

biomedical literature retrieval, large language models, semantic matching, researcher preference modeling, personalization, system architecture, fairness, governance

Abstract

The exponential growth of biomedical literature imposes significant cognitive and operational burdens on researchers who must navigate millions of articles to locate information both specific and serendipitously relevant. Traditional keyword-based search engines, while foundational, often fail to capture the nuanced conceptual interrelations embedded in scientific prose, leading to retrieval results that are poorly aligned with the latent intent of domain experts. This paper presents a comprehensive architectural and systems-level analysis of an AI-enhanced biomedical literature retrieval framework that integrates large language models with a researcher preference-aware semantic matching engine. Rather than proposing a singular algorithm, we examine the design, deployment, governance, and sustainability of a full-stack system that leverages the generative and representational capacities of transformer-based models to construct dynamic user profiles, refine queries through multi-turn interactions, and re-rank candidates using cross-encoder architectures informed by individual research trajectories. We discuss structural trade-offs between latency, precision, and personalization, and investigate how decentralized model fine-tuning, federated preference aggregation, and modular pipeline designs can address robustness and fairness challenges in high-stakes biomedical contexts. The analysis extends to infrastructure considerations, including model compression, edge-cloud orchestration, and energy-aware scheduling, as well as policy dimensions such as data privacy, algorithmic transparency, and the mitigation of confirmation bias that arises from over-personalization. By framing retrieval not as a transactional process but as a socio-technical dialogue between scientist and knowledge base, this paper outlines a principled pathway toward next-generation literature systems that sustain rigorous scientific discovery while remaining ethically grounded and operationally viable.

References

1. Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., & Kang, J. (2020). BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4), 1234-1240.

2. Gu, Y., Tinn, R., Cheng, H., Lucas, M., Usuyama, N., Liu, X., ... & Poon, H. (2021). Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare, 3(1), 1-23.

3. Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., ... & Natarajan, V. (2023). Large language models encode clinical knowledge. Nature, 620(7972), 172-180.

4. Luo, R., Sun, L., Xia, Y., Qin, T., Zhang, S., Poon, H., & Liu, T. Y. (2022). BioGPT: generative pre-trained transformer for biomedical text generation and mining. Briefings in Bioinformatics, 23(6), bbac409.

5. Teevan, J., Dumais, S. T., & Horvitz, E. (2005). Personalizing search via automated analysis of interests and activities. In Proceedings of the 28th annual international ACM SIGIR conference on research and development in information retrieval (pp. 449-456).

6. Huang, P. S., He, X., Gao, J., Deng, L., Acero, A., & Heck, L. (2013). Learning deep structured semantic models for web search using clickthrough data. In Proceedings of the 22nd ACM international conference on information & knowledge management (pp. 2333-2338).

7. Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., ... & Yih, W. T. (2020). Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (pp. 6769-6781).

8. Lopez-Paz, D., & Ranzato, M. A. (2017). Gradient episodic memory for continual learning. In Advances in neural information processing systems (pp. 6467-6476).

9. Zhang, B. H., Lemoine, B., & Mitchell, M. (2018). Mitigating unwanted biases with adversarial learning. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society (pp. 335-340).

10. Yu, X. (2026). Enhancing Search Efficiency through LLM-Based User Memory Systems for Query Matching and Intent Modeling.

11. European Commission. (2021). Proposal for a regulation laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). COM/2021/206 final.

12. Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., ... & Gebru, T. (2019). Model cards for model reporting. In Proceedings of the conference on fairness, accountability, and transparency (pp. 220-229).

13. Roberts, K., Demner-Fushman, D., Voorhees, E. M., Hersh, W. R., Bedrick, S., Lazar, A. J., & Pant, S. (2017). Overview of the TREC 2017 precision medicine track. In The Twenty-Sixth Text REtrieval Conference Proceedings.

14. Topol, E. J. (2019). High-performance medicine: the convergence of human and artificial intelligence. Nature Medicine, 25(1), 44-56.

15. Chen, I. Y., Pierson, E., Rose, S., Joshi, S., Ferryman, K., & Ghassemi, M. (2021). Ethical machine learning in healthcare. Annual Review of Biomedical Data Science, 4, 123-144.

16. Borgman, C. L. (2015). Big data, little data, no data: Scholarship in the networked world. MIT press.

17. Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., ... & Zhao, S. (2021). Advances and open problems in federated learning. Foundations and Trends in Machine Learning, 14(1-2), 1-210.

18. Zuboff, S. (2019). The age of surveillance capitalism: The fight for a human future at the new frontier of power. PublicAffairs.

Downloads

Published

2026-06-30

How to Cite

AI-Enhanced Biomedical Literature Retrieval via Large Language Models and Researcher Preference-Aware Semantic Matching. (2026). Journal of Data Intelligence and AI Systems, 1(3). https://www.jdataai.org/index.php/home/article/view/139