Integrated Clinical, Molecular, and Machine Learning Assessment of Familial Hypercholesterolemia
LIFE-BASEL, vol.16, no.4, 2026 (SCI-Expanded, Scopus)
- Publication Type: Article / Article
- Volume: 16 Issue: 4
- Publication Date: 2026
- Doi Number: 10.3390/life16040633
- Journal Name: LIFE-BASEL
- Journal Indexes: Science Citation Index Expanded (SCI-EXPANDED), Scopus
- Gazi University Affiliated: Yes
Abstract
Background: In clinical practice, LDL-dominant familial hypercholesterolemia (FH) may overlap phenotypically with triglyceride-dominant or mixed familial dyslipidemia. Rule-based diagnostic approaches like the Dutch Lipid Clinic Network (DLCN) and Simon Broome (SB) criteria are frequently used in countries with limited genetic testing, but their concordance with molecular confirmation is inconsistent. In a large Turkish tertiary-care cohort, we studied phenotype-related discordance between clinical criteria and molecular data and tested whether machine learning (ML) models could improve the prediction of reportable pathogenic/likely pathogenic variant positivity among patients with a clinical FH phenotype. Methods: Patients referred for suspected familial hyperlipidemia underwent targeted next-generation sequencing with a 9-gene panel. For the ML analysis, we focused on FH cases with a definitive molecular status (pathogenic/likely pathogenic vs. no reportable variant; variants of uncertain significance were excluded) and applied an 80/20 stratified split (n = 200; 82 molecular-positive cases). Elastic-net logistic regression, random forest, and XGBoost models trained on routinely available clinical variables were compared with dichotomized SB and DLCN classifications. Results: SB positivity was significantly more frequent in triglyceride-dominant phenotypes than in FH (68.4% vs. 52.3%, p = 0.041), despite the substantially lower molecular positivity (14.0% vs. 36.9%, p = 0.002), indicating FH-like false-positive clinical classification in mixed dyslipidemia. In the FH test set, the ML models showed higher discrimination for reportable pathogenic/likely pathogenic variant positivity than dichotomized rule-based criteria (AUC: XGBoost 0.808; random forest 0.769; elastic-net 0.747 vs. SB 0.639; and DLCN 0.598). Thirteen novel variants absent from gnomAD were identified, predominantly in LDLR. Conclusions: In this real-world Turkish cohort, within clinically defined FH cases, ML models performed better at predicting LP/P variant positivity than dichotomized DLCN and Simon Broome criteria. ML-based risk stratification may support prioritization for genetic testing; however, external validation is warranted.