Citation: LY HC, LY TT, NGUYEN DH, et al. Algorithmic consistency and feature-leakage assessment of machine learning models for traditional Chinese medicine syndrome differentiation in herpes zoster: a cross-sectional study. Digital Chinese Medicine, 2026, 9(3): 375-386. DOI: 10.1016/j.dcmed.2026.08.005
Citation: Citation: LY HC, LY TT, NGUYEN DH, et al. Algorithmic consistency and feature-leakage assessment of machine learning models for traditional Chinese medicine syndrome differentiation in herpes zoster: a cross-sectional study. Digital Chinese Medicine, 2026, 9(3): 375-386. DOI: 10.1016/j.dcmed.2026.08.005

Algorithmic consistency and feature-leakage assessment of machine learning models for traditional Chinese medicine syndrome differentiation in herpes zoster: a cross-sectional study

  • Objective To evaluate whether standardized traditional Chinese medicine (TCM) syndrome labels in patients with herpes zoster can be reproduced from structured clinical variables and to determine how model performance changes after removing variables directly used to define the syndrome labels.
    Methods This cross-sectional study included patients with clinically diagnosed herpes zoster at Le Van Thinh Hospital, Ho Chi Minh City, Vietnam, from January 10, 2024 to June 30, 2024. Baseline TCM syndromes were assigned using standardized criteria for liver meridian depression and heat (LMDH) syndrome, spleen deficiency and dampness retention (SDDR) syndrome, and Qi stagnation and blood stasis (QSBS) syndrome. Candidate predictors were audited for feature leakage and categorized as low-, moderate-, or high-risk variables according to their degree of overlap with the criteria used to assign the syndrome labels. Three feature sets were evaluated, including a full-feature set containing all candidate variables, a leakage-reduced feature set excluding direct syndrome-defining variables, and a sensitivity feature set additionally incorporating individual Leeds Assessment of Neuropathic Symptoms and Signs (LANSS) item responses. Full-feature, leakage-reduced, and sensitivity feature sets were evaluated using repeated stratified five-fold cross-validation with five repetitions. Multinomial logistic regression, least absolute shrinkage and selection operator (LASSO)-regularized multinomial logistic regression, decision tree, random forest, support vector machine with a radial basis function kernel, and class-weighted random forest were compared. Macro-F1 score was used as the primary model-selection metric due to class imbalance; balanced accuracy and Cohen’s kappa were reported as secondary evaluation metrics. Overall accuracy was additionally reported as a descriptive metric and interpreted alongside the majority-class no-information rate. Permutation testing was used to determine whether leakage-reduced model performance exceeded that expected after randomization of the syndrome labels.
    Results A total of 80 patients were included in the analysis. LMDH syndrome was the majority class (54/80, 67.5%), followed by SDDR syndrome (16/80, 20.0%) and QSBS syndrome (10/80, 12.5%). In the full-feature analysis, multinomial logistic regression, random forest, and support vector machine achieved an accuracy, macro-F1 score, balanced accuracy, and Cohen’s kappa of 1.000 each, which primarily reflected reproduction of predefined diagnostic rules rather than independent predictive capability. In the leakage-reduced analysis, random forest achieved an accuracy of 0.688, macro-F1 of 0.517, balanced accuracy of 0.494, and Cohen’s kappa of 0.277; class-weighted random forest achieved an accuracy of 0.575, macro-F1 of 0.504, balanced accuracy of 0.545, and Cohen’s kappa of 0.248. Class-specific recall for the leakage-reduced random forest was 0.852 for LMDH, 0.450 for SDDR, and 0.180 for QSBS. Although the accuracy of leakage-reduced random forest model was 0.688, which was higher than the majority-class no-information rate of 0.675, the improvement was limited and balanced multiclass discrimination remained poor. Permutation testing showed that the observed macro-F1 score (0.517) was significantly higher than the mean macro-F1 score (0.356) obtained after the randomized-label permutation (empirical P < 0.001), indicating that some nonrandom classification signal remained after leakage reduction, although balanced three-class discrimination was limited. In the sensitivity analysis feature set, adding individual LANSS item responses did not improve multiclass classification performance.
    Conclusion Full-feature models mainly reproduced the diagnostic rules used to assign TCM syndrome labels. After leakage reduction, the remaining demographic, lesion-related, and pain-related variables showed limited ability to distinguish the three syndrome categories, highlighting the need for feature-leakage auditing and external validation in TCM syndrome machine-learning research.
  • loading

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return