Research Project 03

Health Informatics and Clinical Decision Support

This research develops reliable and interpretable methods for learning from real-world health data. It connects careful clinical study design with patient representation, phenotype discovery, early risk prediction, and evidence that can support clinical decisions.

Integrated health informatics framework from real-world data and cohort design to learning, validation, and clinical decision support
A scalable research framework for turning longitudinal health data into clinically meaningful and validated evidence.

Research philosophy

The clinical question determines the data design and the model. We first define who enters the cohort, what information is available at each time, and what outcome the analysis should estimate. We then choose methods that match the task. These may include clustering, network analysis, prediction, survival modeling, synthetic augmentation, or mixed-methods analysis. Every model must be checked for leakage, calibration, subgroup stability, interpretability, and generalizability before it can inform care.

01

Longitudinal real-world evidence

We build clinically defined cohorts from electronic health records and national data networks. The design fixes the index event, observation period, prediction window, comparison group, and outcome before model training. Matching, weighting, and survival analysis are used when they fit the clinical question.

02

Patient phenotyping and disease trajectories

Patients with the same diagnosis can follow very different clinical paths. Our methods combine patient clustering, variable clustering, and comorbidity networks to identify stable subgroups. The resulting profiles help describe disease progression and reveal patterns that may be missed by a single diagnosis or disease stage.

03

Early prediction with limited clinical data

Many important outcomes are rare, while some treatment groups contain very few patients. We develop prediction models that address imbalance and limited sample size. Synthetic data are used as a training aid within a leakage-controlled evaluation design, not as a replacement for real patients or external validation.

04

Explainable clinical decision support

Prediction alone is not enough for clinical use. We examine which variables influence risk, whether predictions are calibrated, and whether performance remains stable across patient groups and time periods. The final goal is transparent risk stratification that can guide targeted monitoring and follow-up.