带状疱疹中医证候辨证机器学习模型的算法一致性与特征泄漏评估:一项横断面研究

Algorithmic consistency and feature-leakage assessment of machine learning models for traditional Chinese medicine syndrome differentiation in herpes zoster: a cross-sectional study

  • 摘要:
    目的 评价基于结构化临床变量是否能够复现带状疱疹患者标准化中医证候标签,并探讨在去除直接用于定义证候标签的变量后模型性能的变化。
    方法 本研究为横断面研究,纳入2024年1月10日至2024年6月30日在越南胡志明市Le Van Thinh医院临床诊断为带状疱疹的患者。根据标准化中医证候诊断标准,将患者分为肝经郁热证、脾虚湿蕴证和气滞血瘀证。根据候选预测变量与证候判定标准之间的重叠程度,对变量进行特征泄漏风险审查,并将其分为低风险、中风险和高风险变量。研究构建了三种特征集进行评价,包括所有候选变量的完整特征集、排除直接定义证候变量的减少泄漏特征集、加入利兹神经病理性症状和体征评估量表(LANSS)单项条目的敏感性分析特征集。采用重复分层五折交叉验证(重复5次)评估完整特征集、减少泄漏特征集和敏感性分析特征集。比较多项逻辑回归、最小绝对收缩与选择算子(LASSO)正则化多项逻辑回归、决策树、随机森林、基于径向基函数核的支持向量机以及类别加权随机森林模型。由于存在类别不平衡问题,采用宏平均F1值作为主要模型选择指标,平衡准确率和Cohen’s kappa系数作为次要评估指标。总体准确率作为描述性指标,结合多数类别无信息率进行解释。采用置换检验评估减少泄漏模型性能是否超过随机置换证候标签后的预期水平。
    结果 共纳入80例患者进行分析。其中,肝经郁热证为主要类别(54/80,67.5%),其次为脾虚湿蕴证(16/80,20.0%)和气滞血瘀证(10/80,12.5%)。在完整特征集中,多项逻辑回归、随机森林和支持向量机模型实现了准确率、宏平均F1值、平衡准确率和Cohen’s kappa均为1.000的分类性能,该结果主要反映了对预定义诊断规则的复现,而非独立预测能力。在减少泄漏特征集中,随机森林模型的准确率为0.688,宏平均F1值为0.517,平衡准确率为0.494,Cohen’s kappa为0.277;类别加权随机森林模型的准确率为0.575,宏平均F1值为0.504,平衡准确率为0.545,Cohen’s kappa为0.248。减少泄漏随机森林模型对各证候类别的召回率分别为:肝经郁热证0.852、脾虚湿蕴证0.450、气滞血瘀证0.180。尽管减少泄漏随机森林模型的准确率为0.688,略高于多数类别无信息率0.675,但提升幅度有限,且三分类平衡区分能力仍较弱。置换检验显示,观察到的宏平均F1值(0.517)显著高于随机置换标签后获得的宏平均F1值的均值(0.356)(经验性P < 0.001),提示在减少特征泄漏后仍存在部分非随机分类信号,但三类别均衡区分能力仍有限。在敏感性分析特征集中,加入LANSS单项条目并未改善多分类模型性能。
    结论 完整特征模型主要复现了用于分配中医证候标签的诊断规则。在减少特征泄漏后,剩余的人口学特征、皮损相关变量及疼痛相关变量对三类证候的区分能力有限。在中医证候机器学习研究中,应重视特征泄漏审查以及外部验证。

     

    Abstract:
    Objective To evaluate whether standardized traditional Chinese medicine (TCM) syndrome labels in patients with herpes zoster can be reproduced from structured clinical variables and to determine how model performance changes after removing variables directly used to define the syndrome labels.
    Methods This cross-sectional study included patients with clinically diagnosed herpes zoster at Le Van Thinh Hospital, Ho Chi Minh City, Vietnam, from January 10, 2024 to June 30, 2024. Baseline TCM syndromes were assigned using standardized criteria for liver meridian depression and heat (LMDH) syndrome, spleen deficiency and dampness retention (SDDR) syndrome, and Qi stagnation and blood stasis (QSBS) syndrome. Candidate predictors were audited for feature leakage and categorized as low-, moderate-, or high-risk variables according to their degree of overlap with the criteria used to assign the syndrome labels. Three feature sets were evaluated, including a full-feature set containing all candidate variables, a leakage-reduced feature set excluding direct syndrome-defining variables, and a sensitivity feature set additionally incorporating individual Leeds Assessment of Neuropathic Symptoms and Signs (LANSS) item responses. Full-feature, leakage-reduced, and sensitivity feature sets were evaluated using repeated stratified five-fold cross-validation with five repetitions. Multinomial logistic regression, least absolute shrinkage and selection operator (LASSO)-regularized multinomial logistic regression, decision tree, random forest, support vector machine with a radial basis function kernel, and class-weighted random forest were compared. Macro-F1 score was used as the primary model-selection metric due to class imbalance; balanced accuracy and Cohen’s kappa were reported as secondary evaluation metrics. Overall accuracy was additionally reported as a descriptive metric and interpreted alongside the majority-class no-information rate. Permutation testing was used to determine whether leakage-reduced model performance exceeded that expected after randomization of the syndrome labels.
    Results A total of 80 patients were included in the analysis. LMDH syndrome was the majority class (54/80, 67.5%), followed by SDDR syndrome (16/80, 20.0%) and QSBS syndrome (10/80, 12.5%). In the full-feature analysis, multinomial logistic regression, random forest, and support vector machine achieved an accuracy, macro-F1 score, balanced accuracy, and Cohen’s kappa of 1.000 each, which primarily reflected reproduction of predefined diagnostic rules rather than independent predictive capability. In the leakage-reduced analysis, random forest achieved an accuracy of 0.688, macro-F1 of 0.517, balanced accuracy of 0.494, and Cohen’s kappa of 0.277; class-weighted random forest achieved an accuracy of 0.575, macro-F1 of 0.504, balanced accuracy of 0.545, and Cohen’s kappa of 0.248. Class-specific recall for the leakage-reduced random forest was 0.852 for LMDH, 0.450 for SDDR, and 0.180 for QSBS. Although the accuracy of leakage-reduced random forest model was 0.688, which was higher than the majority-class no-information rate of 0.675, the improvement was limited and balanced multiclass discrimination remained poor. Permutation testing showed that the observed macro-F1 score (0.517) was significantly higher than the mean macro-F1 score (0.356) obtained after the randomized-label permutation (empirical P < 0.001), indicating that some nonrandom classification signal remained after leakage reduction, although balanced three-class discrimination was limited. In the sensitivity analysis feature set, adding individual LANSS item responses did not improve multiclass classification performance.
    Conclusion Full-feature models mainly reproduced the diagnostic rules used to assign TCM syndrome labels. After leakage reduction, the remaining demographic, lesion-related, and pain-related variables showed limited ability to distinguish the three syndrome categories, highlighting the need for feature-leakage auditing and external validation in TCM syndrome machine-learning research.

     

/

返回文章
返回