Co-Authored By Varsha Balaji, Canyon Crest Academy Student and SDSC Comms Intern
As large language models (LLMs) become more embedded in healthcare, education, journalism and everyday online search, the stakes of their mistakes grow. Researchers at the University of Illinois Urbana-Champaign (UofI) School of Information Sciences have developed a new way to evaluate whether these systems are producing trustworthy plain-language explanations, helping address a problem that affects how people understand critical information.
LLMs are adept at rewording dense information in a way that can be easily understood by a layperson, but this process, known as plain language summarization (PLS), doesn’t always produce accurate explanations.
To improve the reliability of PLS outputs, researchers Zhiwen You and Yue Guo used U.S. National Science Foundation (NSF) ACCESS allocations on the Delta at the National Center for Supercomputing Applications (NCSA) to develop a means of fact-checking LLMs with an evaluation metric known as PlainQAFact.
The team published their findings in the Journal of Biomedical Informatics.
“We primarily used our NSF ACCESS allocations on NCSA’s Delta to improve the factual consistency evaluation of elaborative explanations that enhance comprehension but are not explicitly present in original scientific sources that the user prompted the LLM to summarize,” said You, a doctoral student at the UofI. “Existing evaluation metrics are capable of fact-checking content when it can be easily compared against the source, while elaborative explanations – even if inaccurate – are overlooked.”
The ability to reduce complex medical knowledge into an easily intelligible summary is crucial to effective communication between scientists in different domains, and by improving the factual consistency of LLMs’ responses to biomedical prompts, we open doors to more accessible and efficient communication.
–Zhiwen You, doctoral student at University of Illinois Urbana-Champaign
The researchers used NCSA’s Delta system to create a way to first sort sentences in a plain language summary as either source simplification or elaborative explanation. Next, the elaborative explanations are isolated and verified using external information from trusted medical sources and a question-and-answer loop between PlainQAFact and the LLM. Finally, PlainQAFact assigns a score to the PLS output, a higher score indicating a more factually consistent response.
“Applied to a biomedical context, PlainQAFact has the potential to meaningfully improve science communications,” You said. “The ability to reduce complex medical knowledge into an easily intelligible summary is crucial to effective communication between scientists in different domains, and by improving the factual consistency of LLMs’ responses to biomedical prompts, we open doors to more accessible and efficient communication.”
Resource Provider Institution(s): National Center for Supercomputing Applications (NCSA)
Resources Used: Delta
Affiliations: University of Illinois Urbana-Champaign
Funding Agency: NSF
Grant or Allocation Number(s): CIS240504
The science story featured here was enabled by the U.S. National Science Foundation’s ACCESS program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.
