This research develops reliable and interpretable methods for learning from real-world health data. It connects careful clinical study design with patient representation, phenotype discovery, early risk prediction, and evidence that can support clinical decisions.
A scalable research framework for turning longitudinal health data into clinically meaningful and validated evidence.
Research philosophy
The clinical question determines the data design and the model. We first define who enters the cohort, what information is available at each time, and what outcome the analysis should estimate. We then choose methods that match the task. These may include clustering, network analysis, prediction, survival modeling, synthetic augmentation, or mixed-methods analysis. Every model must be checked for leakage, calibration, subgroup stability, interpretability, and generalizability before it can inform care.
01
Longitudinal real-world evidence
We build clinically defined cohorts from electronic health records and national data networks. The design fixes the index event, observation period, prediction window, comparison group, and outcome before model training. Matching, weighting, and survival analysis are used when they fit the clinical question.
02
Patient phenotyping and disease trajectories
Patients with the same diagnosis can follow very different clinical paths. Our methods combine patient clustering, variable clustering, and comorbidity networks to identify stable subgroups. The resulting profiles help describe disease progression and reveal patterns that may be missed by a single diagnosis or disease stage.
03
Early prediction with limited clinical data
Many important outcomes are rare, while some treatment groups contain very few patients. We develop prediction models that address imbalance and limited sample size. Synthetic data are used as a training aid within a leakage-controlled evaluation design, not as a replacement for real patients or external validation.
04
Explainable clinical decision support
Prediction alone is not enough for clinical use. We examine which variables influence risk, whether predictions are calibrated, and whether performance remains stable across patient groups and time periods. The final goal is transparent risk stratification that can guide targeted monitoring and follow-up.