Citation: CHEN WL, YE SQ, ZHOU H, et al. TCM-HyperRAG: higher-order knowledge hypergraphs for disease diagnosis and prescription recommendation in traditional Chinese medicine. Digital Chinese Medicine, 2026, 9(3): 362-374. DOI: 10.1016/j.dcmed.2026.08.004
Citation: Citation: CHEN WL, YE SQ, ZHOU H, et al. TCM-HyperRAG: higher-order knowledge hypergraphs for disease diagnosis and prescription recommendation in traditional Chinese medicine. Digital Chinese Medicine, 2026, 9(3): 362-374. DOI: 10.1016/j.dcmed.2026.08.004

TCM-HyperRAG: higher-order knowledge hypergraphs for disease diagnosis and prescription recommendation in traditional Chinese medicine

  • Objective To address the limitations of knowledge graphs in modeling higher-order associations in traditional Chinese medicine (TCM) knowledge and distinguishing the importance of different entities, we propose a weighted knowledge hypergraph-enhanced framework for intelligent assistance in TCM diagnosis and prescription recommendation.
    Methods We developed TCM-HyperRAG, a weighted knowledge hypergraph-based method, and validated it using a publicly available TCM Supervised Fine-Tuning (TCM-SFT) dataset. DeepSeek-V3 extracted disease/syndrome, symptom, and herb entities from 18901 disease-diagnosis records and 9290 prescription-recommendation records to construct the weighted hypergraph. The corresponding retrieval strategy then provided the backbone large language models (LLMs) with structured and traceable evidence. The framework was tested on two held-out sets of 500 previously unseen records for disease diagnosis and prescription recommendation, respectively. LLM-only inference, standard RAG, four GraphRAG modes, and TCM-HyperRAG were compared using HuatuoGPT-o1, Llama3.1-8B, Qwen3-8B, GLM4-9B, DeepSeek-V3, and Qwen3-Max as backbone LLMs. Disease diagnosis was evaluated using mean reciprocal rank (MRR) and hit rate at k (H@k, k ∈ 1, 3, 5), whereas prescription recommendation was evaluated using precision at k (P@k), recall at k (R@k), and F1 score at k (F1@k), where k ∈ 5, 10, 20. Repeated experiments assessed performance stability, and ablation experiments examined the contribution of the importance weights.
    Results TCM-HyperRAG achieved competitive results on both tasks. For disease diagnosis, it obtained the highest MRR with Llama3.1-8B, GLM4-9B, DeepSeek-V3, and Qwen3-Max, and an average MRR of 0.638, higher than 0.571 for LLM-only inference and 0.486 for the GraphRAG variants. For prescription recommendation, TCM-HyperRAG achieved the highest R@20 with all six backbone LLMs, averaging 0.544 versus 0.496 for GraphRAG + Mix. In the repeated experiments with Qwen3-Max, TCM-HyperRAG obtained the highest mean MRR, H@3, and H@5, reaching 0.7691, 0.8625, and 0.9118, respectively, indicating stable performance. Removing entity importance weights reduced the average disease-diagnosis MRR from 0.638 to 0.629 and the average prescription-recommendation F1@10 from 0.438 to 0.425, indicating a positive overall contribution from importance-aware weighting.
    Conclusion By preserving multi-entity relations and prioritizing evidence with greater clinical importance, TCM-HyperRAG achieves competitive performance in disease diagnosis and prescription recommendation while providing traceable evidence for LLM-based intelligent TCM decision support.
  • loading

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return