
==== Front
Perm J
tpj
tpj
The Permanente Journal
1552-5767
1552-5775
The Permanente Press

38952203
10.7812/TPP/24.068
TPJ-24-068
Special Article
editors-choiceEditor’s choiceApplications of Artificial Intelligence in Medicine: An Expert Panel Discussion
http://orcid.org/0000-0003-2305-6755
Halamka John D MD 1
Kirsh Susan R MD, MPH 2
http://orcid.org/0000-0001-6899-9998
Liu Vincent X MD 3
Simon Lynn MD, MBA 4
1 President, Mayo Clinic Platform, Rochester, MI, USA
2 Education and Affiliate Networks, Veterans Health Administration, Washington, DC, USA
3 Kaiser Permanente Division of Research, Santa Clara, CA, USA
4 Community Health Systems, Franklin, TN, USA
John D Halamka, MD halamka.john@mayo.edu
2024
02 7 2024
28 3 312
© 2024 The Authors.
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Published by The Permanente Federation LLC under the terms of the CC BY-NC-ND 4.0 license https://creativecommons.org/licenses/by-nc-nd/4.0/.

Disclaimer The thoughts, ideas, and positions expressed by the panelists in this discussion are their own, do not necessarily reflect the positions of their respective organizations or employers, and do not reflect the position of The Permanente Journal or The Permanente Federation LLC.
==== Body
pmcIntroduction

Since 2022, artificial intelligence (AI) has been the topic of every cocktail party conversation, and yet there are so many unanswered questions as to how we assure the maximization of the benefits of AI within the health care field while minimizing harms. This will require everyone to work together in a public and private partnership.

During this expert panel discussion, we will explore some of the early experiences of the various institutions represented on this panel.

Expert Panel Discussion

John D Halamka

I am John Halamka, and for the last 40 years, I've worked at the intersection of technology and policy. I serve as President of the Mayo Clinic platform overseeing Mayo’s AI initiatives.

Susan R Kirsh

I am Dr Susan Kirsh. I am a general internist and primary care physician with a clinical background.

I've been in the Veterans Health Administration (VHA) for many years, serving as a frontline practitioner and running primary care clinics. For the past 10 years, I have been involved in a number of national efforts, which have included deployment of a clinical contact center and improving health care access with technology, to name a few.

For the past 2 years, I have been the deputy assistant undersecretary for health over the Office of Research and Development, the innovation portfolio, health professions training, and health advancement and partnerships.

Currently, the National Artificial Intelligence Institute reports to the VA Department of Research and Development. As one of the leaders over that office, I have had the opportunity to look at how we operationalize research into practice. This team began under the leadership of Dr Gil Alterovitz. This group has contributed to and is responding to President Biden’s executive orders.

Accordingly, I have learned quite a bit over the last several years and maintain a core focus on how we move new developments into the way we do business to take care of veterans, their families, their caregivers, and the staff. These are the lenses through which I view the topic of this discussion.

Lynn Simon

I appreciate being involved in this discussion. I am Dr Lynn Simon. I'm a neurologist by background and have been at Community Health Systems for about 14 years in a variety of different roles, mostly focused throughout my career on clinical operations. I have experience in enhancing access through a centralized call center within our physician practices. I helped develop a centralized transfer center and physician call center and managed several areas, such as pharmacy, case management, clinical service lines, and our physician enterprise, among others. Around the midpoint of last year, I was asked to take on a role focused on innovation and how we might advance and accelerate our business goals by means such as implementing new services or technologies, leverage partnerships, etc. I look forward to the ensuing conversation.

Vincent X Liu

I'm Vincent Liu. I'm a pulmonary critical care physician by training. I've now been at Kaiser Permanente for 12 years. I am the chief data officer for TPMG (The Permanente Medical Group), which is one of the largest physician-led medical groups in the country. I'm also a research scientist, and a lot of my work focuses on leveraging complex signals from the electronic health record to develop health systems interventions, and to evaluate those interventions to improve patient outcomes and achieve the quadruple aim.

I'm new to the chief data officer role, but there is a lot of activity around AI implementation, responsible AI, AI governance, and augmenting our clinicians with respect to these tools that can allow them to work more effectively while also improving patient care and outcomes.

Biden’s Executive Order on Safe and Effective AI

John D Halamka

Let us move on to some questions to guide our discussion. I am looking for early experiences, given that each of us will have a different perspective on these questions as well as different depths of knowledge. As we look at AI, we all, of course, want fair, appropriate, valid, effective, and safe AI.

We know there will be regulation, legislation, and public–private partnerships. On October 30, 2023, President Biden signed the Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence.1 In that 111-page document, health care is mentioned 33 times. The Executive Order gives a set of what we might consider to be more directives regarding the process to follow as opposed to concrete conclusions.

I would like to open the conversation to the group for initial reaction. Certainly, Dr Kirsh, you have had direct experience with the Executive Order regarding what it says, where it might lead, and some of the early actions we might take.

Susan R Kirsh

Because we are such a large government organization, for us, the Executive Order essentially means that the President and Congress are establishing expectations for our operations. This is the way we do business. For example, the PACT (Promise to Address Comprehensive Toxics) Act, passed by Congress and the VHA, uses new options to increase enrollment for veterans. When we receive executive orders such as this, they are typically filled with high-level policy and operations expectations.

Our organization has had the responsibility of operationalizing some of the expectations. The biggest one right now, in my opinion, is the expectation for trustworthy AI.

We all want AI to be safe and effective. Two things stand out to me most: 1) establishment of trustworthy AI policy and processes, which, as you said, we all want this to be safe and effective, and 2) process operationalizing workforce expectations. Throughout these overarching efforts, data privacy and biases are always part of the consideration.

But how do we extend opportunities for adoption of AI by our frontline clinicians? VA (the US Department of Veterans Affairs) also includes the Veterans Benefit organization and the National Cemetery, so we have a large number of clinicians and staff falling under our purview. How do we help the workers and the leaders understand what this Executive Order means? This has been both an opportunity and a challenge for us. We've set up several cross-agency workgroups and committees. The data governance aspect was already in place, but we are starting to work more effectively across the VA at the secretary level by engaging the chief information officer, the chief technology officer, the clinical leadership at the VHA and the National Artificial Intelligence Institute. We've set up a number of groups that are charged with investigating processes that translate to impact and support for the workers within the organization.

The other big thing I mentioned has been an expectation about where we need to go in developing our workforce. Our secretary has prioritized workforce development and we are working with an interagency group, but within our organization, we are discussing how to upskill people. This is in addition to attracting new data science talent, for which it may be challenging to pay competitive salaries in the federal government.

I am curious to hear what my colleagues in this conversation feel resonates with them or what they might add.

John D Halamka

We know that (President) Biden’s Executive Order included the word “trustworthy.” Now, what is trustworthy AI?

My colleague Casey Ross at STAT, an affiliate publication of The Boston Globe, says AI in health care has a credibility problem. First, it doesn't have transparency, and we don't have measures of consistency or reliability. So, until we have those things, we won't have trustworthiness.

The Executive Order sets up processes (some of which could result in public–private partnerships) to define how we measure the performance of these algorithms. We’ll be talking about different kinds of performance measures next. You may potentially see a nationwide network of AI assurance testing facilities, which would use standards and metrics (the kind of thing we did in the early electronic health record certification era) so that for every algorithm, not only would FDA (the US Food and Drug Administration) oversee safety as it does today, but you might have a nationwide registry that includes a data card (what data are included from various geographies, income/education level, and body types?) and a model card (how does the model perform for all such subgroups?). These types of initiatives would give us a start to approaching trustworthy AI.

But I would argue as well that we need to empower local workforces and not just laboratories. Local workforces must be able to assess whether initiatives will do more harm or good for the patients in front of them. Let me open it up to others for comment.

Lynn Simon

I will begin by reflecting on where we have been and assessing where we’re going in the future. Over the past couple of years, we have evaluated and deployed FDA-approved, but vendor-developed, AI solutions across our portfolio.

Running parallel to that effort, we have brought on our own internal data science team, with large language model (LLM) expertise and deep knowledge and skill sets related to AI, to help us evaluate systems as we go forward. In the meantime, we are starting to build our own platform and develop internal use cases across administrative, clinical, and operational areas.

We also have several partnerships, demonstrating an internal plus external build as opposed to purchasing an existing solution. This depends, of course, on the nature of the individual solutions.

Regarding this evolution and with what Dr Halamka said about trust, many aspects of this were under development before the Executive Order, especially ensuring that our data layer was appropriately brought into an environment in which it could be usable and appropriately harmonized for informing use cases within the organization. This also involved setting up the appropriate internal governance and review structures.

The Executive Order then layered on questions about how the guardrails we've created start to line up with some of these other national guardrails, such as the testing labs that may come from the Executive Order or recommendations from organizations such as CHAI (Coalition for Health AI). This is part of the evolution we will all go through, and, best case, it will help promote more rapid adoption so we can realize the potential of this technology.

I think we have a tremendous opportunity, but we need to make sure that the outcome will project us forward and enable us to solve some of the problems we have in the health care industry today.

Vincent X Liu

We've been focused on the concept of the last mile of implementation and producing results to show that these tools produce benefits for patients.

I can use the illustrative example of the Advance Alert Monitor by way of an example from within TPMG, which we published about in the New England Journal of Medicine.2 This was one of the first examples of AI in that it was a simpler form of machine learning deployed across TPMG’s 21 hospitals. Ultimately, the Advance Alert Monitor produced an estimated benefit of 500 lives saved per year across our hospitals. We have been focusing on that last mile of implementation and the results to show that patients and clinicians can both benefit.

Although we are supportive of a national assurance laboratory framework, I worry that the stamp of approval it provides will be insufficient for what happens in that last mile. I am not so sure whether a model trained and validated in some very prominent health systems or data sources will ultimately still fit the use case or the patient groups for whom they're deployed at the local level. It is important that we also ensure there is adequate training for clinicians through workforce upskilling and augmenting human intelligence to ensure these tools can be deployed at scale in a sustainable, robust, and safe way. That said, I still think there will be a need for rigorous evaluation of these tools and showing proof they work as intended.

I have also had the privilege of leading a program called Kaiser Permanente AIM-HI (Augmented Intelligence in Medicine and Health Care Initiative), a program in which we funded 5 external health systems to deploy AI and machine learning to improve diagnostic decision support. The program extended AI beyond academic or large health system settings and included 2 federally qualified health center networks or safety net institutions. All of the sites will be conducting randomized studies of AI and machine learning in their specific environments. We have to think hard about evaluation and the quality of evidence we need as well as issues with that last mile of integration to arrive at trustworthy AI.

Lynn Simon

I would like to comment on something Dr Liu said regarding sharing initiatives with other health systems. Democratization across health systems, whether they are in academic settings or well-resourced large health systems or small standalone facilities must be part of the consideration. There’s always that question about deploying AI machine learning or generative AI at scale, and how to translate findings and practices across the industry in a way that is feasible for all health systems to take advantage of these technologies. I think it’s an important aspect of the discussion.

Susan R Kirsh

My mind goes to standardization. I came into my career during the quality era with National Quality Forum and The Joint Commission. I think that we inevitably need to head toward a standardization model that has a lot of guardrails around it.

Other organizations that may not have the depth of resources that we have to put into this may be at a disadvantage, and will look to larger organizations like ours. The larger organizations will need to help lead the rollout of initiatives to integrate AI and machine learning into health care, and we should be prepared to provide benchmarks for smaller organizations. There will never be any one perfect metric, and we all have struggled with that, but I think that even with some potential weaknesses, we need to head toward a standard measure to guide us as we lead the way forward.

My organization is involved with the development of an AI assurance model with other government agencies, and we are learning across the government. I feel this effort resembles some of the early quality work groups that came together to put forward the best recommendations.

Related to specific data, we are doing 2 technical sprints per the Executive Order. The first one pertains to ambient dictation software, and the second one pertains to summarizing notes taken from the records of our patients who receive care at health systems outside of the VA. These are fairly common use cases.

We will announce the companies that will win this government challenge at the end of May and will later start with real-time piloting at several sites. The component that resonates the most for me is what the patient thinks about all of this. Of course, cost and efficiency are important as well, but the patient experience also needs to be woven into the evaluation. I think metrics surrounding patient experience and perception need to be mixed into the equation in a prominent manner.

Measuring the Accuracy of Generative AI

John D Halamka

This is a perfect segue to our next question, which is this: As we all begin to deploy generative AI, we have a big question ahead of us. How do we measure the accuracy and the quality of generative AI? With predictive AI, in effect, you can look at an output of a predictive algorithm against ground truth.

We know all algorithms are biased, but bias can be measured. It'll be relatively consistent. For generative AI, a prompt that you put in one minute might yield a brilliant answer, and the same prompt a minute later might yield a fictitious answer.

I welcome the input of the group as to early thinking around measurement of quality and accuracy, or the group may decide that given such a difficult problem, you are limiting the use cases, as Dr Kirsh shared, to more administrative simplification and efficiency gains rather than to anything related to clinical decision support or treatment. I welcome your thoughts on generative AI.

Vincent X Liu

I can kick us off by sharing some of the experiences we had, which have been recently published in the NEJM Catalyst Innovations In Care Delivery.2 The article details some of our early findings with ambient AI. We examined the use of ambient AI in 300,000 clinical encounters among roughly 3000 physicians across TPMG in Northern California. As part of the study, we used an instrument designed to evaluate the quality of human medical scribe documentation to evaluate the quality of the notes taken by the ambient AI scribe tool.

We did some selected sampling of surveys gathered from physicians and patients and also examined more concrete outcomes, like changes in the amount of time users versus nonusers spent in the electronic health record. The results demonstrated that acceptability was high and generally favorable. Overall responses were quite positive. Yet, unsurprisingly, the tools “hallucinate” and they have omissions and other inconsistencies, which we know to be the case with generative AI tools. There are certainly gaps, which relate to their design and statistical grounding, which can poorly integrate context. But largely, we felt that those could be guarded against by ensuring that the physicians were still examining every output.

We found that patient experience was largely improved. Comments such as, “I love talking directly to my doctor without the doctor being behind a computer screen,” were coming through in our results. I find this very promising.

Our research also showed that the time physicians spent writing notes in the electronic health record was reduced. Given the myriad struggles surrounding physician and clinician burnout, anything we can do to reduce the burden of clerical or documentation work, without reducing the quality of those processes, is a big win for the workforce and, in my opinion, is something we should invest in heavily.

Susan R Kirsh

I'll comment briefly that my organization is very risk averse regarding generative AI at this time. Internal policy is that no individual should be working on this across the VHA without explicit support from the chief AI officer. We are not currently ready to purchase products with generative AI features or to test such products in the patient population beyond the customer service–type use cases. We're fairly conservative right now.

Lynn Simon

There is a similar healthy hesitation within most organizations, although perhaps approaching it as a progression that starts with more administrative tasks that are not clinically based and later moving into more clinical operational functions. Certainly, our first foray into embracing AI and machine learning is not going to delve into clinical decision making. If anything, it would simply be more informative for the clinician. I think it’s natural to go through a risk categorization framework, starting with the least risky applications and advancing from there.

That said, we should focus our work to enhance not only the physician experience but also the patient experience, which entails better access to our systems and the ability to get an appointment based around urgency, for instance. It is important to start working in those areas to really change the dynamic of how patients and clinicians interact. Acceptance will be higher when patients and clinicians can see the direct benefit of this technology in reducing the friction we all encounter as health care practitioners or patients.

John D Halamka

What I am hearing from the group is that although there may be a few AI tools out there to support health care, many of us have said our institutions are more likely to select use cases, guardrails, and workflows that assume the underlying generative AI may hallucinate or create fictitious content.

For example, at Mayo, we're using techniques like retrieval-augmented generation. This technology can, for instance, perform functions such as summarizing long medical records, with the user’s understanding that it may have some omissions, but that it’s not likely to hallucinate because it is capable of reasonably summarizing an input. Another example of a use case is to recraft a medical article to the eighth grade reading level if it may be helpful to the patient. In essence, these activities are lower risk.

Deploying AI Tools in Medical Organizations

John D Halamka

Now that we have touched on our impressions of the early benefits of generative AI, it will be interesting to hear from everyone about how your respective organizations have begun to deploy these. Do you have specific areas in which you see the risks and benefits of deploying generative AI in clinical care?

Vincent X Liu

Each of the use cases described highlight the potential benefits, which I think everybody is excited about. Patients’ ability to interact with the health care system is simply far too complex today. The user experience is very different in health care than it is when interacting with the banking system or purchasing things online, for instance. On the whole, it is more challenging for patients to interact with the health care system than it is for these other types of platforms. And although health care is fundamentally different from these other industries, I think there’s hope that tools built using generative AI can reduce the complexity of interacting with health care systems and make it a bit more personal, which in turn can enhance the patient’s care. As for clinicians, they are often buried under a lot of data and the need to address care across several different channels, so there is a lot of excitement about the potential benefits of AI.

Of course, there are also risks. As Dr Halamka described, many of the newest AI tools are changing, so that the same prompt today can produce a very different answer tomorrow for a variety of reasons.

What we see today will also look very different 6 months from now. Generative AI shows a kind of deceptive fluency that can give users the feeling that it understands exactly what’s going on, the context of a query. That’s a risk that we need to be careful about.

Because AI is not human intelligence, we cannot regard it as such. Again, the scale and speed at which it can roll out makes it more challenging to manage. I don't know that we as a health care community have a good solution for that today. I do absolutely agree that the effort will involve patients, health systems, industry partners, researchers, and other public agencies. These various groups will need to come together to decide on the standards, especially in a landscape that is changing incredibly rapidly.

Susan R Kirsh

I agree that we hear a lot about the potential benefits of AI in health care, including about the decrease of administrative burden for various types of clinicians and staff, making it easier for people to do their jobs, which is certainly positive. I had the opportunity to speak in March 2024 on the topic of mental health and AI within the veteran population. We have predictive models using AI and LLMs to evaluate the potential for suicide among our patients that are now being evaluated using AI.

Suicide is an overwhelmingly important clinical problem nationwide, but especially within the Department of Defense and the VHA. We have a program called REACHVET (Recovery Engagement and Coordination for Health—Veterans Enhanced Treatment) is one tool we can use to identify veterans at high risk. We must consider all the potential implications for someone who gets labeled as “high risk.” How do we protect patients who are identified with preexisting conditions using AI algorithms? These are questions I've thought a lot about.

We have to be careful here. Just because you have a risk factor for something doesn't mean that it will definitely happen. Risk can certainly be modified, but modification may or may not be in the model as long as we consider potential risks.

I want to also really promote that we need to learn from each other. I've seen things like CHAI, and Duke runs HAIP (Health AI Partnership), which I've been a little bit involved in, in trying to use academia, public, and private companies and organizations. I think we need to come together on this, and there’s an opportunity there if we do.

Lynn Simon

I agree with what everyone has said. I think that approaching it organizationally requires taking our most important business objectives and outcomes over 1–3 years and determining what tools can help us reach our goals. There must be direct alignment between whatever tools we choose to use with our business objectives. That being said, as an organization starts down this path, it is wise to begin with the tools that are more likely to produce the biggest benefit within the lowest areas of risk.

John D Halamka

Mayo has been working to ensure alignment on planning initiatives across the enterprise. We don't want every individual or department within the organization going off and doing generative AI experimentation on their own.

We decided to coordinate efforts by doing an internal request for applications, indicating that we will make funding available following a competitive process. Ultimately, this brings coordination and resources to these various projects, which would otherwise be disparate. We got 300 applications, which were whittled down to 42, and ultimately to 8 general areas of investigation.

So, we are aligned today regarding various ways in which AI can reduce the burden of administrative tasks, but some of the things we at Mayo are working on are a bit more speculative, none of which are in production but are being explored as interesting, coordinated research efforts.

For example, all of us use genomic sequencing of germ lines and tumors, but how do we interpret mutations and biomarkers into text that can later be turned into action? That’s very difficult. So, we asked the question: “Could you take a million genomes and train a generative AI tool to create text that could offer at least recommendations of other tests to order, lifestyle changes, or other things that could be impactful to the patient?” We have no idea whether this will work, but perhaps it will.

In the notion of multimodal generative AI, what if you took 100 million chest x-rays and 100 million chest x-ray reports? Could generative AI start writing chest x-ray reports? We don't know, but we're going to try.

Let’s look at intake for complex disease states, such as rheumatoid arthritis. It exists as so many different dates with so many different treatment pathways and diagnostic approaches. Could you imagine generative AI taking phenotype, genotype, and exposome and turning it into a care journey? Again, we don't know, but we're trying it out.

One final thing I’ll mention: Maybe medicine is not an LLM, which is predicting the next word in a sentence, but rather is connecting the next event in a care journey.

What if you took the birth-to-death care journeys of 10 million patients, covering all laboratory tests, care encounters, hospitalizations, surgical procedures, weight histories, etc? For a given patient, could you take the benefit of 10 million previous patient journeys and predict the next event that they should have in their health care journey?

Again, we'll try these things, but it’s early.

Let’s shift focus to think about building guardrails and guidelines, such as what Dr Kirsh already mentioned regarding CHAI. There are several public–private partnerships that have also convened. Let’s speculate on the roles of government, academia, and industry in moving us forward toward the credibility and safety that we all seek.

Susan R Kirsh

It’s very creative how Mayo is examining and incentivizing those on the front line. I don't know whether other organizations are doing anything similar. We at the VHA are starting to look toward innovation and research within a testing scenario to support some smaller use cases.

We have the million veteran genome program (Million Veteran Program). Over 10 years, 1 million veterans have enrolled and given blood samples. That group is in research and maybe there is a real opportunity there that you're speaking to, Dr Halamka. It sounds fascinating, and I would love to learn about your outcomes once you are able to share.

John D Halamka

Thank you, Dr Kirsh. You highlighted one of the most important things, as we talk about government, academia, and industry, which is that all of us need to share with each other our early experiences. We should discuss what worked and what didn't work.

Susan R Kirsh

This speaks nicely to government agencies, about harnessing some of the power surrounding what we can do to accelerate progress. There are certain things that government can do well and others where we have challenges. Information technology, privacy, and security are very strict in VA, and we will need to navigate the right balance there.

A considerable amount of effort in scaling initiatives can often be easier in the government than it is in other organizations. We have certain strengths in the government, and we need to look at how we can come together to accelerate and learn, then share findings. Collaboration and learning from each other across various industries and institutions is incredibly important.

John D Halamka

Government can be a wonderful convener and communicator. CHAI now has 2200 organizations, including the FDA, HHS (US Department of Health and Human Services), and OSTP (Office of Science and Technology Policy).

Vincent X Liu

I think ultimately, getting back to this concept of the last mile, the application of these tools within patient care is preserved for the sake of the clinicians who we've trained and who are at the bedside. I think there’s a lot of power in the inferential capability of these statistical models that can get us pretty close to where we need to be. But when it comes to patient care, we need to be better than “pretty close.” I think there will be factors that we can't identify or that generative AI can’t pull together.

Really, when we pair insights generated through these tools with clinicians, I think we're going to produce the maximum benefit. I absolutely think that it requires partnership around the foundational regulatory aspects, the industry, and the way that they're moving in terms of technology, but paired directly with experts at those front lines to understand how we safely integrate these tools into patient care.

With technology moving at its current pace, there will very likely be concerns that the regulations are either too expansive in some cases but may stifle parts of innovation by the time they make their way through to approval due to technology shifts. It will be a dynamic landscape and that will require a lot of communication across these entities.

Susan R Kirsh

I would like to add to what Dr Liu said specifically about FDA. A lot of our clinical staff within our large organization are waiting for FDA to approve these technologies and tools before they can be incorporated into clinical practice. I think there is an opportunity for FDA to develop a model for AI approvals that makes it easier for all of us to use. This includes use of immersive technologies as well. There is an important role for the FDA to play here to approve or support some of these technologies. We will be reliant on them in health care, and the FDA’s stamp of approval is important, although it often does not meet the timelines of these technologies. I am optimistic here in that although the pace is not really keeping up right now, there is an opportunity for FDA to play a big role here.

John D Halamka

I would also suggest that the Office of the National Coordinator for Health Information Technology, with its rules such as HTI-1,3 which requires transparency regarding electronic health record algorithms, will also be an important partner for FDA.

Lynn Simon

Some people have expressed concern around regulatory constructs that would potentially favor the larger incumbents from a technology standpoint. I think we have to be careful that innovators that have a lot to bring to the table are not potentially disadvantaged by what might be unintended constructs limiting their ability to succeed in this very important environment.

Dr Liu’s reference to the last mile speaks to the issue of adoption. As we know, medicine and the health care industry have not historically been exceptionally good at rapid adoption of new techniques and technologies. So, how do we get these coalitions to help us work together to speed up adoption, which likely ties back in with trust and transparency?

There are so many problems to solve, and I'm excited that AI and generative AI are some very promising technologies that have come along recently to show great promise. And yet, how do we integrate that promise into the way we provide care and conduct our business more quickly than we have in the past? I believe this question warrants further consideration.

John D Halamka

That’s well said. We should emphasize to our readers that this is not about hegemony of any one organization. Rather, it is about empowering all organizations, not simply the large academic medical centers. It needs to be the community hospitals and the rural clinics too, for example. These also need to be participants in the measurement and validation of our algorithms, but also in the deployment and adoption of them. We must all work diligently to ensure that happens.

AI in Medicine: Hype Versus Reality

John D Halamka

My last question is regarding AI hype versus reality.

I was talking to a Silicon Valley leader who shall remain unnamed. This person claims that generative AI is already sentient. This leader went further to suggest that we don't need to worry about curating large datasets for training because the generative AI will simply read all digital data that exist and teach itself, and soon physicians will no longer be necessary because, according to this leader, all physicians do is pattern match. The sentiment offered was that generative AI, which is sentient and has learned more than any single physician, can do a better job of pattern matching.

So, I will share that in my personal view of the hype versus reality, I would strongly suggest that all of these aforementioned assertions made by said Silicon Valley leader are false. Let me turn it over to the group for comment.

Lynn Simon

I am concerned about this kind of rhetoric. I worry that talking points like this will dominate the conversation and prevent us from having the conversations that actually matter. If replacement theory dialogue takes front and center, we aren’t focusing on the possibility of really leveraging AI in health care to augment and enhance the performance of our professionals.

John D Halamka

Very well said.

Vincent X Liu

We should absolutely be very careful about humanizing AI. We have human intelligence, and we have artificial intelligence. What we've seen is that AI, at least according to standardized tests and other interactive tools, actually can achieve what we have traditionally viewed as human intelligence. But I think to anthropomorphize AI by comparing it to ourselves as humans actually does a disservice to AI because the best future state is to let the machines do what they do best and incorporate their strengths with human intelligence. We should understand that we may exhibit similar accuracies, but where we humans falter will be different from where the machine falters. In actuality, by working together, there’s an opportunity to create tremendous enhancement.

It also highlights a lack of understanding about the complexity of health care from some of our industry partners. So, we have to simply continue to be engaged and lead awareness about these complexities. That being said, technology innovations have disrupted many industries even when that very same industry’s leaders have claimed to be far too complex to be fed into an app, for instance. There is a real concern about what the future will look like. The only way we can make the future what we want is to stay engaged and continue to lead.

Susan R Kirsh

I echo what my colleagues on this panel have said. As we learned in the pandemic, being isolated and interfacing only virtually for certain things can only get you so far. You still need human interaction to live a happy and fulfilling life. This is true for most of us. We can be human, and let the machines be machines.

Defining and clarifying boundaries is evolving and important. As an optimist, I think there’s an opportunity, but we are going to have to do all the things discussed on this call, such as standing up the partnerships and setting boundaries in a way that is prudent for clinical care. I still think that humans need to be part of the interaction augmented by our technologies and tools.

John D Halamka

The reality of generative AI is that emergent properties don't suddenly appear; they evolve gradually. As humans, we don't go from "no math" to "all math." We make mistakes as we learn addition and multiplication. Generative AI has the same behavior as parameters are increased: no math, some math with limited accuracy, better accuracy. Actually, as we did more and more parameters, it was a linear progression of capabilities and not an emergent property that was an artifact of what we were measuring. So be careful what you read and how you interpret it.

I would like to hear final thoughts from each of the panelists, please.

Lynn Simon

I simply echo the optimism that has been described in this panel. We made it through COVID-19. It was a very disruptive time that left us with many problems within health care to solve. But at the same time, we are starting to see emerging tools and techniques that can be used to solve some of those problems.

I regard the emergence of AI within health care as having the strong potential to solve challenges such as those that were created over the last several years. I remain optimistic that we will address these challenges. There is just enough of a burning platform in the health care industry that will be the impetus to push us in the direction to do just that.

Susan R Kirsh

As for 2024, we are eagerly awaiting the completion of the tech sprints, which will serve as our first national foray into learning what the impact of AI tools in our organization might be. We are investigating integration into our own workflows and are taking into consideration lessons learned as a government entity. We will be continuing to work with the Office of the National Coordinator for Health Information Technology, FDA, HHS, industry, and academia to continue to learn and contribute in any way that we can.

We want to enable learnings at the executive level while also providing learning and resources for the frontline practitioners and teams. To me, in my role as a leader in this organization, providing these resources to frontline workers is probably one of the most important things I want to focus on. Frontline staff are intimately familiar with the biggest problems within the organization. Therefore, enabling, facilitating, and supporting our frontline folks to address challenges is inspiring. So, I appreciate you calling that out specifically.

Vincent X Liu

I would say that all of us play the role of data stewards on behalf of our organizations and our patients. Our roles are about preserving the privacy, security, and trustworthiness of the data, but also not letting the data go to waste, thereby preventing us from generating the insights into the tools that we need. I think a lot of our focus is on maximizing the trustworthy use of these data while also maximizing the preservation of privacy and security. That’s a challenge, but it’s an exciting time to be working with data and caring for patients.

John D Halamka

Thank you to all of the panelists who joined to discuss AI in health care. I will leave the readers of The Permanente Journal with the thought that AI will not replace physicians, nurses, or other clinicians, but that the health care workforce can use AI to augment their workflows, and those who do will outperform those who don't.

Conflicts of Interest: None declared

Funding: None declared
==== Refs
References

1. The White House . Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence. October 30, 2023. Accessed 4 May 2024. https://www.whitehouse.gov/briefing-room/presidential-actions/2023/10/30/executive-order-on-the-safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence/
2. Tierney AA , Gayre G , Hoberman B , et al. Ambient artificial intelligence scribes to alleviate the burden of clinical documentation. NEJM Catalyst. 2024;5 (3 ). 10.1056/CAT.23.0404
3. The Office of the National Coordinator for Health Information Technology . Health data, technology, and interoperability: Certification program updates, algorithm transparency, and information sharing (HTI-1) final rule. Accessed 4 May 2024. https://www.healthit.gov/topic/laws-regulation-and-policy/health-data-technology-and-interoperability-certification-program
