TCM-HyperRAG:面向中医疾病诊断和处方推荐的高阶知识超图

TCM-HyperRAG: higher-order knowledge hypergraphs for disease diagnosis and prescription recommendation in traditional Chinese medicine

  • 摘要:
    目的 针对知识图谱难以建模中医药知识中的高阶关联,且无法区分不同实体重要性的问题,本研究提出一种加权知识超图增强框架,用于中医智能辅助诊断和处方推荐。
    方法 本研究构建了基于加权知识超图的方法TCM-HyperRAG,并基于公开的中医药监督微调数据集(TCM-SFT)进行验证。DeepSeek-V3逐条从18 901条疾病诊断记录和9 290条处方推荐记录中抽取疾病或证候、症状和中药实体,构建加权知识超图,并通过相应的检索策略为骨干大型语言模型(LLM)提供结构化且可追溯的证据。在疾病诊断和处方推荐任务上分别采用500条未见病例进行测试,以HuatuoGPT-o1、Llama3.1-8B、Qwen3-8B、GLM4-9B、DeepSeek-V3和Qwen3-Max为骨干LLM,对单纯LLM推理、标准RAG、4种GraphRAG模式和TCM-HyperRAG进行比较。疾病诊断采用平均倒数排名(MRR)和前 k 项命中率(H@ k , k \in \1,\;3,\;5 \)评价,处方推荐采用前 k 项准确率(P@k)、前 k 项召回率(R@k)和前 k 项F1分数(F1@k)( k \in \5,\;10,\;20\ )评价。通过重复实验评价模型表现的稳定性,并通过消融实验验证重要性权重的贡献。
    结果 TCM-HyperRAG在两个任务上均取得了具有竞争力的结果。在疾病诊断任务中,TCM-HyperRAG在Llama3.1-8B、GLM4-9B、DeepSeek-V3和Qwen3-Max取得最高MRR,平均MRR为0.638,高于纯LLM推理的0.571和GraphRAG变体的0.486。在处方推荐任务中,TCM-HyperRAG在全部6种大模型上均取得最高R@20,平均值为0.544,高于GraphRAG+Mix的0.496。在基于Qwen3-Max的重复实验中,TCM-HyperRAG取得最高的平均MRR、H@3和H@5,分别为0.76910.86250.9118,表明该方法具有稳定的表现。移除实体重要性权重后,疾病诊断平均MRR由0.638降至0.629,处方推荐平均F1@10由0.438降至0.425,表明权重感知对整体性能具有积极贡献。
    结论 TCM-HyperRAG通过保留多实体关系并优先检索临床重要性较高的证据,在疾病诊断和处方推荐任务中取得了具有竞争力的结果,同时可为基于LLM的中医智能决策支持提供可追溯的证据。

     

    Abstract:
    Objective To address the limitations of knowledge graphs in modeling higher-order associations in traditional Chinese medicine (TCM) knowledge and distinguishing the importance of different entities, we propose a weighted knowledge hypergraph-enhanced framework for intelligent assistance in TCM diagnosis and prescription recommendation.
    Methods We developed TCM-HyperRAG, a weighted knowledge hypergraph-based method, and validated it using a publicly available TCM Supervised Fine-Tuning (TCM-SFT) dataset. DeepSeek-V3 extracted disease/syndrome, symptom, and herb entities from 18901 disease-diagnosis records and 9290 prescription-recommendation records to construct the weighted hypergraph. The corresponding retrieval strategy then provided the backbone large language models (LLMs) with structured and traceable evidence. The framework was tested on two held-out sets of 500 previously unseen records for disease diagnosis and prescription recommendation, respectively. LLM-only inference, standard RAG, four GraphRAG modes, and TCM-HyperRAG were compared using HuatuoGPT-o1, Llama3.1-8B, Qwen3-8B, GLM4-9B, DeepSeek-V3, and Qwen3-Max as backbone LLMs. Disease diagnosis was evaluated using mean reciprocal rank (MRR) and hit rate at k (H@k, k ∈ 1, 3, 5), whereas prescription recommendation was evaluated using precision at k (P@k), recall at k (R@k), and F1 score at k (F1@k), where k ∈ 5, 10, 20. Repeated experiments assessed performance stability, and ablation experiments examined the contribution of the importance weights.
    Results TCM-HyperRAG achieved competitive results on both tasks. For disease diagnosis, it obtained the highest MRR with Llama3.1-8B, GLM4-9B, DeepSeek-V3, and Qwen3-Max, and an average MRR of 0.638, higher than 0.571 for LLM-only inference and 0.486 for the GraphRAG variants. For prescription recommendation, TCM-HyperRAG achieved the highest R@20 with all six backbone LLMs, averaging 0.544 versus 0.496 for GraphRAG + Mix. In the repeated experiments with Qwen3-Max, TCM-HyperRAG obtained the highest mean MRR, H@3, and H@5, reaching 0.7691, 0.8625, and 0.9118, respectively, indicating stable performance. Removing entity importance weights reduced the average disease-diagnosis MRR from 0.638 to 0.629 and the average prescription-recommendation F1@10 from 0.438 to 0.425, indicating a positive overall contribution from importance-aware weighting.
    Conclusion By preserving multi-entity relations and prioritizing evidence with greater clinical importance, TCM-HyperRAG achieves competitive performance in disease diagnosis and prescription recommendation while providing traceable evidence for LLM-based intelligent TCM decision support.

     

/

返回文章
返回