
==== Front
Perm J
tpj
tpj
The Permanente Journal
1552-5767
1552-5775
The Permanente Press

39113535
10.7812/TPP/24.064
TPJ-24-064
Commentary
editors-choiceEditor’s choiceGaining Trust: Lessons and Opportunities for Artificial Intelligence in Health Care
http://orcid.org/0000-0002-9377-1714
Wallace Paul J MD 1
1 Retired. Past affiliations with Kaiser Permanente, United Health Group, and AcademyHealth, OR, USA
Paul J Wallace, MD Paul.Wallace@sbcglobal.net
2024
08 8 2024
28 3 168171
© 2024 The Authors.
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Published by The Permanente Federation LLC under the terms of the CC BY-NC-ND 4.0 license https://creativecommons.org/licenses/by-nc-nd/4.0/.

Keywords:

Trustworthiness
Artificial Intelligence
Evidence-based Medicine
==== Body
pmcIn this issue of The Permanente Journal, an expert panel addresses the use of artificial intelligence (AI) in health care.1 Drs Halamka, Kirsh, Liu, and Simon have done an admirable job, noting the varied range of applications under the AI umbrella, including large language models/generative AI, machine learning, natural language processing, and others. They speak as leaders within organizations traditionally at the cutting edge in care design and delivery. Their discussion features the pursuit of appropriate consistency, including public and private investments in national assurance facilities, juxtaposed with respect for the essential role of the “final mile” where end-user clinicians and patients are critical to both applying and improving the emerging tools to fit local requirements. They also recognize the challenges of democratization, to ensure that the benefits of AI are not restricted to the most privileged organizations and open to a range of innovators.

The discussion details how they are systematically testing organizational strategies for AI deployment. Although the approaches differ, common tactics include: careful analysis of proposed use cases; active pursuit of wide support by leaders and those impacted by potential change; meaningful engagement with innovators supported by careful observation of results and impact; and the intent to share learnings about both successes and failures within and outside of their organizations. Further, the excitement over AI is reinforced by the range and potential effects of their initiatives. These include assessment of the impact of ambient scribes on the work and care experience of clinicians and patients and testing predictive models and deep learning to address challenging situations like the danger of suicide among higher-risk populations or the development of sepsis in hospitalized patients. They also note early investments to highly personalize care ranging from use of genomics to systematic characterization of patients with complex diseases, and finally, testing if new care insights can emerge from studying the “care journey” over time of increasingly complex individuals.

The overall discussion reflects optimism, the value of collaboration, and the potential for AI to support clinicians and patients in making effective judgments. That said, AI is a new model for generating and communicating knowledge. The valid use of AI in health care decision-making is emerging but also largely unproven. There are substantial concerns about its use for specific patients and situations as well as the presence of conflicts of interest and distortion of findings due to identifiable bias.

I strongly agree with the roundtable contributors that AI tools are best positioned to complement the agency and autonomy of health care decision-makers, not to replace or minimize their roles. However, to fulfill this promise, advocates for AI have an urgent responsibility to demonstrate that guidance derived using AI tools can be trusted.

Trust and Trustworthiness

Decision-making in health care is a special case and crafting support for choices that can impact well-being, wealth, and survival itself requires exceptional diligence to earn trust. Fortunately, there are robust precedents for gauging trustworthiness and building trust in health care knowledge creation and sharing paradigms.

For over a century, exploration and learning in medicine and health care have primarily built upon disciplined and serial hypothesis testing. However, not all results generated by traditional discovery have proven equally trustworthy. Reliability of findings can be compromised by knowable factors, including: inadequate study design resulting in bias or insufficient power to justify an apparent finding, the failure to report and reconcile conflicts of interest, and an inability to create results that translate directly into meaningful action by end users.

The model that has emerged to systematically address this challenge for traditional health science is evidence-based medicine (EBM). “Evidence based medicine is the conscientious, explicit, and judicious use of current best evidence in making decisions about the care of individual patients.”2 Notably, EBM is built on work since the 1960s, largely by librarians as early data scientists, to digitize and hence improve the ability to methodically search, vet, and organize the health-related scientific literature around a central concern or question.

EBM has been widely accepted as the standard of trustworthiness across organizations and at the front lines of care. Clinical practice guidelines (CPG), developed to guide care for large numbers of patients, are a use of EBM. Similar efforts are embedded in new technology assessment, which I discussed in a commentary in The Permanente Journal in 2001.3 EBM principles also underlie the quality improvement movement, including work to improve patient safety as well as widely implemented systems of performance measurement. A further evolution of evidence-based practice has been the systematic approach to improving population health, especially for people affected by chronic health conditions such as diabetes, heart disease, and asthma.

Systematically Demonstrating Trustworthiness

Applications of EBM have been robustly challenged by skeptical stakeholders. CPGs offer an in-depth example of persistent testing, refinement, and advocacy over decades to earn trust by patients and care practitioners. This quest culminated in a pair of congressionally mandated reports from the Institute of Medicine (IOM) in 2011, Finding What Works in Health Care: Standards for Systematic Reviews 4 and Clinical Practice Guidelines We Can Trust.5

This effort proposed standards that “reflect a review of the literature, public comment, and expert consensus on best practices for developing trustworthy guidelines” emphasizing transparency, management of conflict of interest, establishing evidence foundations for guideline recommendations, plus requirements for external review and updating. Through similar scrutiny, EBM-related applications have been answerable for ensuring fidelity with sound science and best care while still promoting progress and access to emerging innovation.

Toward a Framework for Trustworthy AI

Again, advances in AI represent new streams for knowledge development and communication. Testing traditional health science against the standards for EBM ensures that findings are trustworthy and appropriately scalable for high-impact circumstances. For health care applications that rely on AI to be similarly seen as trustworthy, developers and advocates can follow a similar path to proactively earn accountability.

The demonstration of trustworthiness is multifactorial. The durable framework for CPGs is an integrated sum of multiple components driven and informed by public and expert stakeholders. All of the elements are required. Developing a similar approach for AI and its application to important health decisions is urgent. To seed that work, components of the structure identified in the IOM reports may be adapted and refined as best to apply to AI.

Possible key elements include at least the following:

Transparency. This is the “who, what, when, where, and how” of application development that should be consistently disclosed by developers, advocates, and end users to enable stakeholders to understand how recommendations were derived and who developed them.

Identification and Management of Conflict of Interest. Conflicts of interest are “a set of circumstances that creates a risk that professional judgment or actions regarding a primary interest will be unduly influenced by a secondary interest.”6 Conflicts that should be disclosed can include organizational, financial, personal, and relational interests that may influence the results of an AI application.

Identification and Attention to Sources of Bias. The potential for a finding to be distorted due to methodologic choices exists in all approaches to discovery. Bias exists in many forms,7 with common concerns including information bias, selection bias, and confounding. Developers and advocates should discern potential biases and note how risk is minimized by the approach chosen. For example, a major benefit of randomized controlled trials for pursuit of traditional clinical discovery is to minimize selection bias.

External Review. For CPGs, the IOM report recommends that “external reviewers should comprise a full spectrum of relevant stakeholders, including scientific and clinical experts, organizations (eg, health care, specialty societies), agencies (eg, federal government), patients, and representatives of the public.” A similar expectation for at least high-impact AI applications may also address the following concerns:

Is oversight beyond the marketplace appropriate for some or all applications?

Are there requirements for peer review of findings? Is publication in peer-reviewed journals an expectation?

Which applications should be reviewed by regulatory bodies such as the Food and Drug Administration or the proposed Assurance Testing Facilities?

Is a central clearinghouse for health-related AI applications required

Updating; When do AI applications outdate? Should there be a schedule for review, revalidation, or planned senescence?

Validation: Corroboration of a causal link between the data used and a proposed action is the core tension encountered in integrating many AI findings within care recommendations. The systematic translation of EBM-validated findings into practice provides an analog. However, justifying the use of novel AI-derived discoveries as an appropriate and trustworthy basis for clinical guidance will require substantial new diligence. Concerns to be addressed include:

How are patterns in data sought, recognized, and brought forward?

How is an identified correlation connected to a meaningful aspect of causation?

What is the strength of that association? How compelling is the association?

How can findings be validated? Is a secondary trial needed for validation? In what circumstances?

Forming a Trusted Partnership

The standards above are the entry requirements for AI to earn a role in guiding health decisions. Arguably, any AI application should provide users with real-time responses to all of the standards. As noted in the roundtable,1 AI also has the potential to provide insights that have not emerged from traditional health science. Some examples:

An already-manifested concern is that AI will be employed to amplify historical approaches to utilization review of insurance claims and used to restrict insurance coverage. Alternatively, as individuals are identified as outliers from expected parameters of care, the ability to look across large populations can surface care strategies that have been successful in highly similar situations and allow a pivot from coverage sanction to promotion of an appropriately personalized intervention.

Patients with multiple chronic conditions are common but poorly addressed by existing research and care guidelines, which generally address a single condition. Consequently, the care for these complex individuals can be highly variable within and among practices and frustrating for both clinicians and patients. AI presents the ability to look across large populations to dynamically assemble cohorts that closely resemble a given complex patient in the moment and over time and then, based on experiences elsewhere, identify effective care strategies as guidance for that highly selective circumstance.

The management of a busy practice entails large volumes of data review, analysis, decision-making, and documentation. Testing how best to leverage and integrate AI-supported information flow and management has huge potential. Prioritizing the prevention of clinician burnout complemented by applications that foster patient participation and aspects of selfcare would be timely.

Positioning AI capabilities as trusted partners for addressing existing clinician and patient workflow and decision-making challenges would be expected to pay high dividends.

Conclusion

As exemplars of innovation, AI tools have enormous potential. AI use can be expected to complement and extend existing health care knowledge and care roles, while also offering solutions to historically difficult problems. However, for that to occur, the ultimate but surmountable test will be for AI applications and their advocates to earn our trust.

Author Contribution: Paul Wallace, MD, conceptualized, drafted, and submitted the final manuscript.

Conflicts of Interest: None declared

Funding: None declared
==== Refs
References

1. Halamka JD , Kirsh SR , Liu VX , Simon L . Applications of Artificial Intelligence in Medicine: An Expert Panel Discussion. Perm J. 2024;28 (3 ):3–12. 10.7812/TPP/24.068
2. Sackett DL , Rosenberg WM , Gray JA , Haynes RB , Richardson WS . Evidence based medicine: What it is and what it isn’t. BMJ. 1996;312 (7023 ):71–72. 10.1136/bmj.312.7023.71 8555924
3. Wallace P . Addressing the challenge of new medical technologies: One Permanente clinician’s view – part 1. Perm J. 2001;5 (2 ):72–74.
4. Institute of Medicine . Finding What Works in Health Care: Standards for Systematic Reviews. The National Academies Press; 2011. 10.17226/13059
5. Institute of Medicine . Clinical Practice Guidelines We Can Trust. The National Academies Press; 2011. 10.17226/13058
6. Brems JH , Davis AE , Clayton EW . Analysis of conflict of interest policies among organizations producing clinical practice guidelines. PLOS One. 2021;16 (4 ). 10.1371/journal.pone.0249267
7. Berkman ND , Santaguida PL , Viswanathan M , et al. The empirical evidence of bias in trials measuring treatment differences [Rockville, MD: Agency for Healthcare Research and Quality (US]. Accessed 28 07 2014. https://www.ncbi.nlm.nih.gov/books/NBK253181/
