Grounded Factual Language Generation for Hallucination Mitigation in Large Language Models
A research plan for evidence-grounded, transparent and explainable language generation, with particular attention to Spanish.
Complete bibliography
References
18 references cited in the paper
- J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. GPT-4 Technical Report. Technical Report, OpenAI, 2024.
- M. Cossio. A comprehensive taxonomy of hallucinations in large language models. arXiv preprint, 2025.
- L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, T. Liu. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems 43, 2025, 1–55.
- A. J. Oche, A. G. Folashade, T. Ghosal, A. Biswas. A systematic review of key retrieval-augmented generation (RAG) systems: Progress, gaps, and future directions. arXiv preprint, 2025.
- Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, M. Wang, H. Wang. Retrieval-augmented generation for large language models: A survey. arXiv preprint, 2024.
- H. Liu, Z. Wang, X. Chen, Z. Li, F. Xiong, Q. Yu, W. Zhang. HopRAG: Multi-hop reasoning for logic-aware retrieval-augmented generation. arXiv preprint, 2025.
- T. Ni, X. Yuan, S. Li, K. Wu, R. Liu, W. Ni, W. Zhang. Stepchain GraphRAG: Reasoning over knowledge graphs for multi-hop question answering. arXiv preprint, 2025.
- Q. Zhang, Z. Xiang, Y. Xiao, L. Wang, J. Li, X. Wang, J. Su. FaithfulRAG: Fact-level conflict modeling for context-faithful retrieval-augmented generation. arXiv preprint, 2025.
- A. P. Alodjants, A. E. Avdyushina, D. V. Tsarev, I. A. Bessmertny, A. Y. Khrennikov. Quantum approach for contextual search, retrieval, and ranking of classical information. Entropy 26, 2024, 862.
- X. Li, R. Zhao, Y. K. Chia, B. Ding, S. Joty, S. Poria, L. Bing. Chain-of-knowledge: Grounding large language models via dynamic knowledge adapting over heterogeneous sources. arXiv preprint, 2024.
- S. Dhuliawala, M. Komeili, J. Xu, R. Raileanu, X. Li, A. Celikyilmaz, J. Weston. Chain-of-verification reduces hallucination in large language models. arXiv preprint, 2023.
- P. Sen, A. F. Aji, A. Saffari. Mintaka: A complex, natural, and multilingual dataset for end-to-end question answering. Proceedings of COLING 2022, pp. 1604–1619.
- P. Lewis, B. Oğuz, R. Rinott, S. Riedel, H. Schwenk. MLQA: Evaluating cross-lingual extractive question answering. Proceedings of ACL 2020, pp. 7315–7330.
- M. Artetxe, S. Ruder, D. Yogatama. On the cross-lingual transferability of monolingual representations. Proceedings of ACL 2020, pp. 4623–4637.
- J. Li, X. Cheng, W. X. Zhao, J.-Y. Nie, J.-R. Wen. HaluEval: A large-scale hallucination evaluation benchmark for large language models. Proceedings of EMNLP 2023, pp. 6449–6464.
- D. Edge, H. Trinh, N. Cheng, J. Bradley, A. Chao, A. Mody, S. Truitt, J. Larson. From local to global: A graph RAG approach to query-focused summarization. arXiv preprint, 2024.
- A. Asai, Z. Wu, Y. Wang, A. Sil, H. Hajishirzi. Self-RAG: Learning to retrieve, generate, and critique through self-reflection. ICLR, 2024.
- S. Min, K. Krishna, X. Lyu, M. Lewis, W.-t. Yih, P. W. Koh, M. Iyyer, L. Zettlemoyer, H. Hajishirzi. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. Proceedings of EMNLP 2023, pp. 12076–12100.