Grounded Factual Language Generation for Hallucination Mitigation in Large Language Models

A research plan for evidence-grounded, transparent and explainable language generation, with particular attention to Spanish.

References

18 references cited in the paper

  1. J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. GPT-4 Technical Report. Technical Report, OpenAI, 2024.
  2. M. Cossio. A comprehensive taxonomy of hallucinations in large language models. arXiv preprint, 2025.
  3. L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, T. Liu. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems 43, 2025, 1–55.
  4. A. J. Oche, A. G. Folashade, T. Ghosal, A. Biswas. A systematic review of key retrieval-augmented generation (RAG) systems: Progress, gaps, and future directions. arXiv preprint, 2025.
  5. Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, M. Wang, H. Wang. Retrieval-augmented generation for large language models: A survey. arXiv preprint, 2024.
  6. H. Liu, Z. Wang, X. Chen, Z. Li, F. Xiong, Q. Yu, W. Zhang. HopRAG: Multi-hop reasoning for logic-aware retrieval-augmented generation. arXiv preprint, 2025.
  7. T. Ni, X. Yuan, S. Li, K. Wu, R. Liu, W. Ni, W. Zhang. Stepchain GraphRAG: Reasoning over knowledge graphs for multi-hop question answering. arXiv preprint, 2025.
  8. Q. Zhang, Z. Xiang, Y. Xiao, L. Wang, J. Li, X. Wang, J. Su. FaithfulRAG: Fact-level conflict modeling for context-faithful retrieval-augmented generation. arXiv preprint, 2025.
  9. A. P. Alodjants, A. E. Avdyushina, D. V. Tsarev, I. A. Bessmertny, A. Y. Khrennikov. Quantum approach for contextual search, retrieval, and ranking of classical information. Entropy 26, 2024, 862.
  10. X. Li, R. Zhao, Y. K. Chia, B. Ding, S. Joty, S. Poria, L. Bing. Chain-of-knowledge: Grounding large language models via dynamic knowledge adapting over heterogeneous sources. arXiv preprint, 2024.
  11. S. Dhuliawala, M. Komeili, J. Xu, R. Raileanu, X. Li, A. Celikyilmaz, J. Weston. Chain-of-verification reduces hallucination in large language models. arXiv preprint, 2023.
  12. P. Sen, A. F. Aji, A. Saffari. Mintaka: A complex, natural, and multilingual dataset for end-to-end question answering. Proceedings of COLING 2022, pp. 1604–1619.
  13. P. Lewis, B. Oğuz, R. Rinott, S. Riedel, H. Schwenk. MLQA: Evaluating cross-lingual extractive question answering. Proceedings of ACL 2020, pp. 7315–7330.
  14. M. Artetxe, S. Ruder, D. Yogatama. On the cross-lingual transferability of monolingual representations. Proceedings of ACL 2020, pp. 4623–4637.
  15. J. Li, X. Cheng, W. X. Zhao, J.-Y. Nie, J.-R. Wen. HaluEval: A large-scale hallucination evaluation benchmark for large language models. Proceedings of EMNLP 2023, pp. 6449–6464.
  16. D. Edge, H. Trinh, N. Cheng, J. Bradley, A. Chao, A. Mody, S. Truitt, J. Larson. From local to global: A graph RAG approach to query-focused summarization. arXiv preprint, 2024.
  17. A. Asai, Z. Wu, Y. Wang, A. Sil, H. Hajishirzi. Self-RAG: Learning to retrieve, generate, and critique through self-reflection. ICLR, 2024.
  18. S. Min, K. Krishna, X. Lyu, M. Lewis, W.-t. Yih, P. W. Koh, M. Iyyer, L. Zettlemoyer, H. Hajishirzi. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. Proceedings of EMNLP 2023, pp. 12076–12100.