高级检索

大语言模型可解释性研究综述

A Survey on the Explainability of Large Language Model

  • 摘要: 大语言模型凭借出色的任务解决能力而备受关注。从基础语言建模和文本生成任务到复杂推理任务,大模型实现了从通用能力到专业能力的转变,并在与用户的交互中逐步落地于各应用场景。当前,尽管大模型已产生深远的影响,但是仍然因为内部机制透明度低、伦理道德水平待考量等可解释问题而备受质疑,需要更多针对大模型的可解释性研究以揭示其内部机理。由于大模型的输出受到交互过程及应用领域的规范的影响,文中从模型和用户这2个更加立体的角度,将大模型的可解释性拆分成模型决策过程透明度、用户交互可控性和模型输出结果可信度,对现有工作进行了全面综述。首先从大模型本身出发,对模型内、外部紧扣训练微调技术和增强技术分类介绍现有的解释方法;然后从人机交互层面介绍从用户端输入提示引导模型决策,增强大模型可解释性的研究;最后指出大模型可解释性研究面临的缺乏可解释性的精准定义、公认的评估标准等局限性。

     

    Abstract: Large language models have gained prominence due to their outstanding task-solving capabilities. The evolution from basic language modeling and text generation tasks to complex reasoning tasks has facilitated the transition of large language models from general to specialized capabilities. This gradual implementation across various application scenarios underscores their utility in interaction with users. Despite the unprecedented and profound impact of these models, they are often criticized for their lack of transparency in internal mechanisms and ethical considerations. There remains a necessity for more explainability research to unveil the mystery of large language models, thereby enhancing their abilities to adapt to downstream tasks and improve user experience. The outputs of large language models are influenced by both the interactive process and the strict norms of domain applications. This paper innovatively proposes a comprehensive review of explainability studies of large language models from two dimensions: model and user.Specifically, we combine model process-transparency and interaction-controllability with the credibility of model outputs. First, we explore the model itself, categorizing existing interpretative methods based on internal and external techniques for training fine-tuning and enhancement. Next, from the perspective of human-computer interaction, we discuss research on guiding model decisions through user-input prompts to enhance model explainability. Finally, we outline the limitations faced by explainability studies of large language models, such as the lack of a precise definition of interpretability and universally recognized evaluation standards.

     

/

返回文章
返回