them-pure

AI in Nephrology: Why the Hardest Part Was Never the Algorithm

Artificial intelligence has quietly become one of the most active areas of research in nephrology. Across acute kidney injury, chronic kidney disease, dialysis care, and kidney transplantation, machine learning and deep learning models are now routinely reported with strong predictive performance, in many cases rivalling or exceeding the accuracy of the risk tools nephrologists have relied on for decades. Yet a pattern keeps repeating across this body of research: models that perform impressively in a paper rarely make it into the room where a clinician is actually deciding what to do next. Understanding why that gap exists, and what it will take to close it, matters just as much as understanding what these models can do.

A Field Moving in Four Directions at Once

Recent state-of-the-art reviews of AI in nephrology, including a comprehensive 2026 roadmap published in Clinical Kidney Journal, describe a field advancing simultaneously across acute injury, chronic disease, dialysis, and transplantation, each with a distinct set of use cases.

In acute kidney injury, predictive models trained on high-frequency electronic health record data and intensive care unit telemetry have shown strong performance in forecasting the onset of AKI before it becomes clinically obvious, giving care teams a window to intervene earlier than traditional serum creatinine monitoring allows. In chronic dialysis care, machine learning models are increasingly used to guide anemia management and to forecast intradialytic hypotension, a common and clinically significant complication of hemodialysis.

Models using recurrent neural networks and gradient-boosted trees have demonstrated the ability to anticipate these hypotensive episodes up to sixty minutes in advance, with accuracy scores frequently exceeding 0.85, and this predictive lead time has been associated with reductions in treatment interruptions and hospitalisation rates when acted upon. In transplantation, AI applications are being explored for donor-recipient matching and post-transplant outcome prediction, while generative AI and large language models are beginning to find a role in clinical documentation, patient triage, and health education, tasks that sit adjacent to direct clinical decision-making but still meaningfully reduce the administrative load on care teams.

Taken together, this represents genuine, well-documented progress. But it is in chronic kidney disease specifically, the area with arguably the largest population impact given how common and how frequently under-detected CKD is, where the promise and the practical challenges of AI in nephrology come into sharpest focus.

Where AI Is Reshaping CKD Prediction

Chronic kidney disease has always been difficult to manage well partly because of a basic limitation in how risk has traditionally been assessed. Conventional tools such as the CKD Epidemiology Collaboration equation and the Kidney Failure Risk Equation, while clinically useful and widely adopted, are limited in their precision, their generalisability across different populations, and, critically, their ability to identify which patients in the earliest stages of CKD are likely to progress rapidly toward kidney failure versus those who will remain stable for years.

This is precisely the gap recent AI research has been targeting. A 2026 review in the International Urology and Nephrology journal examined how machine learning, deep learning, natural language processing, and multimodal data integration are being applied to improve CKD detection, progression prediction, and personalised management.

Drawing on data from electronic health records, medical imaging, omics profiling, and increasingly wearable devices, these AI approaches have achieved accuracy scores, measured by area under the curve, ranging from 0.85 to 0.96 when predicting outcomes such as progression to end-stage kidney disease and a patient's likely response to a given treatment.

For context, an AUC of 1.0 would represent perfect prediction, and scores above 0.85 are generally considered strong performance in clinical risk modelling. Beyond simple risk scores, these models are also being used for phenotype clustering, grouping patients by shared underlying disease patterns rather than a single risk number, which opens the door to more individualised surveillance schedules and treatment strategies rather than a one-size-fits-all approach to CKD management.

This matters enormously for a disease that is so often diagnosed late. A model capable of flagging, early and accurately, which patients are likely to be rapid progressors could fundamentally change how CKD is managed, shifting care from a reactive response to declining lab values toward a genuinely proactive surveillance strategy built around each patient's actual trajectory.

The Gap Between the Paper and the Clinic

And yet, despite these consistently strong accuracy figures across dozens of published models, the honest state of the field is that very few of these tools are actually being used in day-to-day clinical practice. This disconnect is not a minor footnote in the research. It is, increasingly, the central conversation happening within nephrology's own literature.

Several recurring barriers explain why. The same Clinical Kidney Journal review that catalogued AI's promise in dialysis care noted plainly that few AI applications in nephrology have undergone formal evaluation of their cost-effectiveness, their feasibility for insurance reimbursement, or their long-term sustainability once deployed outside a research setting. Implementation itself carries real costs, in data infrastructure, in integration with existing electronic health record systems, and in the ongoing technical maintenance a live clinical model requires, costs that a strong AUC score in a published paper does not account for.

A parallel review focused specifically on intradialytic hypotension prediction observed that AI dashboards can provide real-time alerts and support prevention, but limited validation and inconsistent methodology across different studies have impeded broader adoption, even where the underlying predictive accuracy looks genuinely promising on paper.

The problem runs deeper than infrastructure and cost alone. A systematic review and meta-analysis of AI models predicting CKD prognosis found that while these tools achieved a strong pooled accuracy score of 0.89 overall, they struggled to balance sensitivity and specificity well. The models were very good, with a specificity of 0.92, at correctly identifying patients who would not progress, but comparatively weak, with a sensitivity of only 0.43, at reliably catching the patients who would.

In plain terms, a model can report an excellent-looking headline accuracy number while still missing more than half of the patients it most needs to catch early, a distinction that matters enormously in clinical practice but can be easy to overlook in a summary statistic. The American Society of Nephrology's own 2026 statement on responsible AI use in kidney care named this cluster of issues directly, pointing to data quality, ethical concerns, the sheer complexity of integrating a new tool into existing clinical workflows, and, perhaps most fundamentally, clinician trust, as the barriers standing between promising research and real bedside use.

This is, in many ways, the defining tension of AI in nephrology today. The algorithms themselves are, by most published measures, already good. What has consistently lagged is the harder, less glamorous work of making these tools trustworthy, interpretable, and genuinely usable within the rhythm of an actual clinical encounter, rather than existing as a standalone research artefact that a clinician would need to consult separately, interpret independently, and reconcile against their own judgement without support.

What It Will Actually Take to Close That Gap

Closing this gap is not simply a matter of publishing more models with even higher accuracy scores. Based on where the field's own research and professional guidance is converging, a few directions stand out as genuinely necessary rather than merely aspirational.

The first is building models as workflow-native tools rather than standalone outputs. A recent framework published in the Journal of Clinical Medicine describes this shift explicitly, arguing that the next generation of clinical AI in nephrology needs to move from prediction toward action, functioning as systems that continuously perceive patient data, reason within real clinical constraints, and support coordinated next steps directly within the setting where care decisions are actually made, rather than requiring a clinician to leave their workflow to consult a separate dashboard.

The second is keeping the clinician firmly in the loop rather than positioning AI as an autonomous decision-maker. The American Society of Nephrology's AI Workgroup built this principle into the foundation of its guidance, stating explicitly that physician oversight must remain a non-negotiable assumption underlying any responsible clinical AI tool, with the technology positioned to inform and accelerate a clinician's judgement rather than to replace it.

The third is prioritising real-world, prospective validation across genuinely diverse patient populations and care settings, rather than relying solely on the retrospective, single-centre datasets that produce the impressively high but often fragile accuracy figures seen in early-stage research. And the fourth is transparency in how a model actually arrives at its output, using interpretable techniques that show clinicians which specific factors are driving a given risk score, since a recommendation a clinician cannot understand or interrogate is a recommendation they are, quite reasonably, unlikely to trust with a patient's care.

Designing for the Point of Care, Not Around It

This is precisely the thinking behind the CKD staging heatmap we are building for the Proflo-U® app platform, and it is why we have approached it less as a predictive model in isolation and more as a decision-support layer designed to sit exactly where the clinical action already happens.

Once complete, the system is designed to take the relevant kidney health parameters captured at the point of testing, which are uACR and eGFR, and automatically map them onto the internationally recognised KDIGO albuminuria-GFR risk matrix. Rather than surfacing a raw risk score that still requires separate clinical interpretation, the output is intended to be an accurate, colour-coded risk classification paired with a structured recommendation: continue routine annual screening, initiate targeted intervention, retest in three months, or refer urgently to nephrology. Crucially, this output is designed to be generated at the same moment and in the same setting as the test itself, rather than requiring a clinician to later reconcile a separate analytics report against their own judgement.

Just as importantly, the system is not designed to make the final call on its own. Every classification and recommendation it generates is intended to be presented as exactly that, a recommendation, put up for clinical review and referral rather than an autonomous verdict. The operator or clinician using the platform retains the ability to weigh that recommendation against everything else they know about the patient in front of them and make the precise decision themselves. This is a deliberate design choice, not an incidental one. It reflects the same principle nephrology's own professional guidance has converged on: that AI earns clinical trust not by replacing judgement, but by giving it better, faster, more consistently structured information to work with, at the exact moment that judgement needs to be exercised.

Why This Distinction Matters More Than It Might Seem

It would be easy to read the difference between "an AI model that predicts CKD risk" and "an AI-supported staging tool embedded directly into a point-of-care testing workflow, kept under clinical review" as a small implementation detail. It is not. It is, based on where the field's own literature keeps pointing, the actual difference between a tool that exists convincingly on paper and one that has a realistic chance of being used, consistently, by a health worker or clinician who does not have the time, training, or inclination to interpret a standalone risk score in isolation.

The models with genuinely strong predictive accuracy already exist across AKI, CKD, and dialysis care. What nephrology has been missing is not better mathematics. It is tools built with enough humility about how clinical work actually happens, tools that show up inside the workflow rather than beside it, that keep a clinician's judgement at the centre rather than at the periphery, and that are honest about being decision support rather than decision replacement. Building toward that, deliberately and carefully, is what we believe gives AI in nephrology its best chance of moving from an impressive body of research into something that actually changes outcomes for patients.

Sources: Cheungpasitporn W, et al. Transforming nephrology through artificial intelligence: a state-of-the-art roadmap for clinical integration. Clinical Kidney Journal, 2026;19(2):sfag004. Yuan S, et al. Artificial intelligence in nephrology: predicting CKD progression and personalizing treatment. International Urology and Nephrology, 2026;58(7):2615-2645. Tangri N, Cheungpasitporn W, et al. Responsible Use of Artificial Intelligence to Improve Kidney Care: A Statement from the American Society of Nephrology. Journal of the American Society of Nephrology, 2026;37(4):881-890.



Social Share