Abstract
Clinical artificial intelligence (AI) systems are increasingly used to support healthcare decision-making, yet standard evaluation approaches often assume that electronic health record (EHR) observations are comparable across patients and care contexts. This assumption is problematic in settings where recorded data are produced through complex, variable, and resource-constrained care processes. This dissertation examines that problem in in-center hemodialysis (ICHD), a high-frequency care environment where treatment depends on repeated attendance, completed dialysis sessions, routine monitoring, and ongoing medication adjustment. In an ideal setting, care would occur as prescribed. In practice, prescribed care and delivered care may diverge because of interacting factors, such as facility policy and transportation availability, within the sociotechnical system. This dissertation conceptualizes that divergence as data distortion and focuses specifically on care instability, defined as recurring disruptions in care delivery, as one source of this distortion shaping the EHR data used to train, evaluate, and interpret clinical AI systems.
Using a mixed methods design, this dissertation develops and applies a context-stratified framework for evaluating fairness in an AI system designed for treating chronic kidney disease-mineral and bone disorder (CKD-MBD) in ICHD. First, a scoping review examines existing AI and machine learning research on treating CKD-MBD in ICHD, and how fairness, bias, and sociotechnical context were included. Second, the dissertation develops the Vulnerability Index (Vulnerability Index), a dynamic, individual-level measure of care instability in ICHD derived from routinely collected treatment records, including missed and shortened dialysis sessions. Third, the Vulnerability Index is used to evaluate whether clinical AI behavior for CKD-MBD treatment recommendations varies across categories of care instability and care conditions. Fourth, qualitative interviews with dialysis patients and health professionals contextualize the mechanisms that produce care instability and identify aspects of dialysis care that may not be fully visible in structured EHR data and thus may impact AI fairness throughout ICHD.
Findings show that existing AI research on CKD-MBD has largely emphasized prediction, recommendation, and model performance while giving limited attention to fairness-relevant care contexts. The Vulnerability Index provides a reproducible way to operationalize care instability from routine ICHD treatment records. When applied to a CKD-MBD treatment recommender AI system, results show that clinical AI behavior can vary across Vulnerability Index categories in ways that aggregate evaluation may obscure. Qualitative findings further demonstrate that care instability reflects more than individual treatment participation, emerging instead from interacting societal and structural constraints affecting access to care.
This dissertation contributes a systems engineering approach to clinical AI evaluation under care instability in ICHD. By treating care instability as a measurable feature of the clinical data-generating process, this work shows how context-stratified evaluation can reveal fairness-relevant variation in AI behavior that may be missed by aggregate performance metrics alone.