
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

72528
10.1038/s41598-024-72528-3
Article
The Two Word Test as a semantic benchmark for large language models
Riccardi Nicholas 1
Yang Xuan 2
Desai Rutvik H. rutvik@sc.edu

2
1 https://ror.org/02b6qw903 grid.254567.7 0000 0000 9075 106X Department of Communication Sciences and Disorders, University of South Carolina, Columbia, 29208 USA
2 https://ror.org/02b6qw903 grid.254567.7 0000 0000 9075 106X Department of Psychology, University of South Carolina, Columbia, 29208 USA
16 9 2024
16 9 2024
2024
14 215937 5 2024
9 9 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Large language models (LLMs) have shown remarkable abilities recently, including passing advanced professional exams and demanding benchmark tests. This performance has led many to suggest that they are close to achieving humanlike or “true” understanding of language, and even artificial general intelligence (AGI). Here, we provide a new open-source benchmark, the Two Word Test (TWT), that can assess semantic abilities of LLMs using two-word phrases in a task that can be performed relatively easily by humans without advanced training. Combining multiple words into a single concept is a fundamental linguistic and conceptual operation routinely performed by people. The test requires meaningfulness judgments of 1768 noun-noun combinations that have been rated as meaningful (e.g., baby boy) or as having low meaningfulness (e.g., goat sky) by human raters. This novel test differs from existing benchmarks that rely on logical reasoning, inference, puzzle-solving, or domain expertise. We provide versions of the task that probe meaningfulness ratings on a 0–4 scale as well as binary judgments. With both versions, we conducted a series of experiments using the TWT on GPT-4, GPT-3.5, Claude-3-Optus, and Gemini-1-Pro-001. Results demonstrated that, compared to humans, all models performed relatively poorly at rating meaningfulness of these phrases. GPT-3.5-turbo, Gemini-1.0-Pro-001 and GPT-4-turbo were also unable to make binary discriminations between sensible and nonsense phrases, with these models consistently judging nonsensical phrases as making sense. Claude-3-Opus made a substantial improvement in binary discrimination of combinatorial phrases but was still significantly worse than human performance. The TWT can be used to understand and assess the limitations of current LLMs, and potentially improve them. The test also reminds us that caution is warranted in attributing “true” or human-level understanding to LLMs based only on tests that are challenging for humans.

Subject terms

Psychology
Computer science
http://dx.doi.org/10.13039/100000002 National Institutes of Health 5R01DC017162-02 5R01DC017162-02 5R01DC017162-02 Riccardi Nicholas Yang Xuan Desai Rutvik H. issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Large Language Models (LLMs; also called Large Pre-Trained Models or Foundation Models1) are deep neural networks with billions or trillions of parameters that are trained on massive natural language corpora. They have shown surprising and remarkable abilities spanning many different tasks. Some examples include the ability to pass examinations required for advanced degrees, such as those in law2, business3, and medicine4. Strong performance on benchmarks such as General Language Understanding Evaluation (GLUE) and its successor (SuperGLUE) have also been obtained5,6. Bubeck et al.7 investigated an early version of GPT-4, and reported that it can solve difficult tasks in mathematics, coding, vision, medicine, law, psychology, and music, and exhibited “mastery of language.” With such breadth of human-level (or better) performance, they suggested that it shows “sparks” of Artificial General Intelligence (AGI).

Such achievements have led many researchers to conclude that LLMs have achieved or are close to achieving  a real or humanlike understanding of language. Others remain skeptical. A recent survey8 asked active researchers whether such models, trained only on text, could in principle understand natural language someday. About half (51%) agreed, while the other half (49%) disagreed. This stark divide is closely tied to the question of what constitutes true understanding and has been the subject of intense debate9.

The skeptics have pointed out examples where LLMs produce less-than-satisfactory performance. Hallucinations10,11, inaccurate number comparisons, and reasoning errors are commonly cited problems, and failures in individual cases are frequently reported (e.g., https://github.com/giuven95/chatgpt-failures). It is argued that while LLMs exhibit formal linguistic competence, they lack functional linguistic competence, which is the ability to robustly understand and use language in the real world12. However, this claim still runs into the problem of how to measure robust understanding beyond subjective assessments of the quality of answers in response to prompts. Objective benchmarks are essential here, but as successes and failures of LLMs show, benchmarks that are suitable for measuring human understanding might not be appropriate for assessing LLMs13–15.

There are philosophical arguments as to why LLMs do not have true or humanlike understanding. For example, LLMs learn words-to-words mappings, but not words-to-world mappings, and hence cannot understand the objects or events that words refer to16. Such arguments aside, formal tests are critical, as that’s where “rubber meets the road.” If a system can match or surpass human performance in any task thrown at it, the argument that it does not possess real understanding rings hollow. If an LLM indeed lacks humanlike understanding, one ought to be able to design tests where it performs worse than humans. With such tests, the nebulous definition of “understanding” becomes less of a problem.

Here, we propose and evaluate one such novel benchmark, the Two Word Test (TWT). The test is based on a basic human psycholinguistic ability to understand combinations of two words, which has been of great interest to psychologists, linguists, and neuroscientists for decades17–20. The test uses noun-noun combinations such as beach ball, and requires discrimination between meaningful and nonsense or low meaningfulness combinations. Compared to other linguistic possibilities, such as adjective-noun (big ball) or verb-noun (throw ball) combinations, noun-noun combinations do not offer grammatical assistance in determining meaningfulness. One can determine that ball red is not a meaningful phrase, because noun-adjective is not a valid word order in English. The same strategy cannot be used to determine that ball beach has low meaningfulness. Some phrases are learned as single units that combine unrelated words (sea lion), while others are “built from the ground up”. Baby boy makes sense, and many other words could follow baby and the phrase would still be sensible (clothes, girl, sister, etc.). Simply reversing the word order of some of these (clothes baby) can result in a low-meaningfulness phrase. While the exact cognitive processes involved in human word combination remain debated20, it requires the composition of two or more lexical elements into a coherent whole, followed by a judgment of that whole constituent against prior world knowledge and related concepts. In this context, phrase-level familiarity/frequency is a vital part of ‘meaningfulness’, but does not fully explain it. Phrases that we encounter more often will naturally be accepted as meaningful, but some low-frequency phrases can also be highly meaningful. Unlike existing benchmarks that rely on the ability to do logical reasoning, planning, infer implied content, disambiguate ambiguous words, or solve puzzles or other problems, the TWT has unique memory-dependent, semantic, and compositional elements that make it a valuable semantic benchmark for LLMs. A previous study21 obtained meaningfulness ratings on these phrases from 150 human participants. In the current study, we aim to use comprehensive statistical tests to compare the judgment of humans and four current LLMs (OpenAI’s GPT-4-turbo and GPT-3.5-turbo, Google’s Gemini-1-Pro, and Anthropic's Claude-3-Opus) in the original TWT measuring the meaningfulness judgment using a Likert scale, and a binary TWT (bTWT) measuring binary “makes sense” or “nonsense” judgments. Finally, we examine the effect of n-gram phrase frequency and semantic cosine similarity between the two words in the presented phrases on human and LLM meaningfulness ratings to identify possible reasons for discrepancies between LLM and human ratings. We discuss the limitation of current LLMs in language comprehension reflected in these novel benchmarks as a weakness distinct from those in tasks that rely on higher-level executive control, such as logical reasoning or puzzle solving.

Materials and methods

Two word test phrase generation and human rating collection

The TWT consists of noun-noun combinations and human meaningfulness ratings collected as part of behavioral and neuroimaging experiments conducted by Graves and colleagues21, whose methods we will now summarize. They chose 500 common nouns, and all possible noun-noun combinations were generated, resulting in 249,500 potential phrases. The occurrence of these combinations as two-word phrases was cross-referenced with a large corpus of human-generated text, and phrases that appeared at least once and in only one direction (i.e., ‘noun1 noun2’ but not ‘noun2 noun1’) were kept. Then, as judged by Graves et al., phrases with meaningful interchangeable word orders or that were taboo were removed, resulting in 1080 possibly meaningful phrases. Possible “nonsense” or low-meaningfulness phrases were then generated by reversing the word order of meaningful phrases, resulting in 2,160 total phrases for rating by humans (half being possibly ‘meaningful’ and half being possibly ‘nonsense’). Note that, at this point, no actual rating by humans had been conducted yet, so it is possible that some of the possible ‘meaningful’ phrases would be given low meaningfulness ratings by humans, or vice versa. Participants (N = 150), who were undergraduates at the University of Wisconsin-Madison in the early 2010s, then rated subsets of the total phrase pool with the following instructions:

Please read each phrase, then judge how meaningful it is as a single concept, using a scale from 0 to 4 as follows: If the phrase makes no sense, the appropriate rating is 0. If the phrase makes some sense, the appropriate rating is 2. If the phrase makes complete sense, the appropriate rating is a 4. Please consider the full range of the scale when making your ratings. Examples: the goat sky, 0 (makes no sense), the fox mask, 2 (makes some sense), and the computer programmer, 4 (makes complete sense).

For each phrase, the mean and standard deviation of participant responses were calculated. 392 phrases with mean ratings between 1.5 and 2.5 were removed from the set due to being ambiguous to human raters. This resulted in 977 nonsense and 761 meaningful phrases, as judged by human raters, resulting in the final set of 1,768 word pairs used in the two word test (TWT) presented here. 81.7% of these pairs appear twice in the TWT stimuli, once forward and once backward. The remainder of pairs indicate instances where human raters judged a word pair as either meaningful or nonsense, but that word pair’s reverse was judged as ambiguous (e.g., rated between 1.5 and 2.5). The final set of stimuli can be found at https://github.com/NickRiccardi/two-word-test.

This dataset was chosen for this benchmark for a few reasons. First, it is one of the largest available sets of combinatorial noun-noun phrases that have corresponding human ratings. Second, due to how the list was constructed, it provides good control of important psycholinguistic variables such as word frequency and concreteness that may have profound effects on human and LLM ratings. That is, because low-sensibility phrases are created by reversing the word order of high-sensibility phrases, many lexical confounds related to psycholinguistic properties between the low- and high-sensibility phrases are eliminated.

Results and discussion

We conducted a series of experiments comparing Claude-3-Opus, Gemini-1.0-Pro-001, GPT-4-turbo, and GPT-3.5-turbo performance (each model as available in March 2024) to the human data. First, we gave the LLMs the same prompt used by Graves et al.21, followed by an enhanced prompt. Then, we tested the LLMs on a binary version of the test (i.e., “makes sense”/“nonsense” judgment instead of numerical ratings) that was expected to be easier for LLMs.

TWT: numerical meaningfulness judgments

Due to token restrictions, the phrases were randomly assigned to 8 subsets to ensure that the LLMs' errors were not due to memory limitations. To prompt the LLMs' judgment, we submitted the same instructions and examples originally provided by Graves et al. However, using Graves' original prompt resulted in the LLMs largely neglecting the use of the 1 and 3 ratings, the two ratings not used as example cases in Graves' original prompt. To encourage the LLMs to use the full rating scale, we provided two additional examples in the instructions for scores of 1 and 3 (the knife army, 1 (makes very little sense), and the soap bubble, 3 (makes a lot of sense)). Each query input consisted of the original instructions, four examples, and a subset of the phrases appended to the end. To rule out the possibility that the LLMs' meaningfulness judgment on two-word phrases depended on the meaningfulness of the phrases prior to or following it, we shuffled the order of the phrases ten times and repeated the query for ten iterations. The query inputs were kept the same for different LLMs to ensure a direct comparison. The final meaningfulness judgment for each phrase was the average score across ten iterations. The 1768 unambiguous phrases were selected. Compared to the human distribution, which reflects “makes sense” and “nonsense” phrases in the bimodal peaks, Gemini-1.0-Pro and GPT-3.5-turbo showed a bias towards rating most phrases as a 2 or 3 (makes some sense, makes a lot of sense; Fig. 1). GPT-4-turbo and Claude-3-Opus showed similar bimodal peaks as humans, but there were still a considerable number of phrases that were judged as ambiguous with a rating around 2 (makes some sense). See supplementary materials for scatterplots comparing human ratings of each item to the ratings of LLMs.Fig. 1 Frequency of continuous meaningfulness ratings for humans and LLMs. Human mean responses reflect a bimodal distribution of meaningful and nonsense phrases, while that is lacking in all four LLMs.

However, it is more informative to take LLM ratings of each individual phrase and test the probability that its rating came from the same distribution as the human responses to that phrase. We conducted a series of phrase-wise statistical tests to compare each LLM to human meaningfulness ratings.

First, we used the human means and standard deviations for each individual phrase (provided by Graves et al.21 to generate a Gaussian distribution of 10,000 simulated human responses to each phrase, respecting the lower and upper limits of the 0-to-4 scale, resulting in an approximation of realistic response distributions (Fig. 2; see supplementary materials for additional examples). Then, for each phrase, we conducted a Crawford and Howell t-test22 for case–control comparisons with the LLM as the case and the human distribution as the control. This modified t-test is designed for the comparison of a single-case observation to a control group and returns the probability that the case comes from the same distribution as the group. We hereby define a “TWT failure” as when the LLM meaningfulness rating has less than a 5% probability of coming from the human distribution (i.e., the LLM rating is significantly different from that of humans, p < 0.05).Fig. 2 Simulated human rating distributions (gray) and LLM ratings for low- and high-meaningfulness phrases (the cake apple, the dog sled). For the cake apple, GPT-3.5-turbo, GPT-4-turbo, and Gemini-1.0-Pro-001 rated it as more meaningful than > 95% of humans would be expected to, while Claude-3-Opus responded within normal limits. For the dog sled, GPT-4-turbo rated it as less meaningful than > 95% of humans would be expected to, while the other LLMs responded within normal limits.

To understand where LLM responses fall between human ratings and chance or random ratings, we generated two rating distributions. (1) “Human”: 1000 simulated participants whose phrase-wise responses were generated from the underlying probability distribution of responses to each phrase in Graves et al.21. (2) “Chance”: 1000 permuted participants whose phrase-wise responses were selected based only on the frequency of 0–4 ratings from the original study. The “Human” distribution approximates what would be expected from human raters if the study was run on a large number of human participants. The “Chance” distribution is what would be generated by a system with no knowledge of word meaning. We then generated failure counts for the distributions and for each of the models.

Experiment 1 Results: Table 1 and Fig. 3 show that Gemini-1.0-Pro, GPT-3.5-turbo, and GPT-4-turbo failure counts are closer to chance than to the simulated human distribution. Claude-3-Opus is significantly better than the other LLMs, but still fails far more than what would be expected from a human rater. Taken together, these results show that the three LLMs fail at the TWT, but that there are significant differences between their abilities.Table 1 Summary of TWT failures.

LLM	Failure counts	Failure percentages	
claude-3-Opus	331	18.7	
gemini-1.0-Pro-001	648	36.7	
gpt-3.5-turbo	974	55.1	
gpt-4-turbo	515	29.1	

Fig. 3 Number of LLM failures in TWT compared to simulated human (blue) and permuted-chance (orange) failure count distributions.

bTWT: binary Two Word Test

LLMs are often reported to make errors on numerical tasks. It is possible that the poor performance on the TWT was due to a difficulty in dealing with the numerical scale required for the task, rather than a lack of understanding of phrase meaning. In order to eliminate numerical ratings, we modified the TWT instructions to prompt binary responses, again appending phrases to the end of the instructions in randomized order and with multiple submissions:Please read each phrase, then judge how meaningful it is as a single concept. If the phrase makes no sense or makes very little sense, the appropriate response is “nonsense”. If the phrase makes a lot of sense or complete sense, the appropriate rating is “makes sense”. Examples: “the goat sky” is “nonsense”, “the knife army” is “nonsense”, “the soap bubble” is “makes sense”, “the computer programmer” is “makes sense”.

For comparison, we recoded the numerical ratings of human judgments to binary values. Out of the 1768 phrases, 977 phrases (mean ratings larger than 2.5) were recoded as “make sense” and 761 phrases (mean ratings less than 1.5) were recoded as “nonsense”. We then calculated the following to measure LLM performance: Chi-squared (χ2) test, signal detection theory (SDT) metrics, and receiver operating characteristic (ROC) curve.

Experiment 2 Results: Table 2 and Fig. 4 show SDT results. SDT measures how well an actor (LLMs) can detect true signals (meaningful phrases) while correctly rejecting noise (nonsense phrases). It uses ratios of hits (true positives), correct rejections (true negatives), false alarms (false positives), and misses (false negatives). d’ is a measure of overall ability to discriminate, with 0 being chance-level and > 4 being very high discrimination ability. β measures response tendency, or whether an actor prefers to say that a signal is present (liberal) or absent (conservative). The base 10 logarithm of β, reported here, is interpreted as < 0 being liberal and > 0 being conservative. We also display the ROC curve (Fig. 5) and report the area under the curve (AUC).Table 2 d', β, and AUC for LLM “makes sense” / “nonsense” discrimination.

LLM	d’	β	AUC	
claude-3-Opus	1.85	− 0.37	0.80	
gemini-1.0-Pro-001	1.66	− 0.50	0.75	
gpt-3.5-turbo	0.35	− 0.47	0.50	
gpt-4-turbo	1.33	− 0.28	0.72	

Fig. 4 SDT metrics for all 1,768 phrases. Hit – true positive; Miss – false negative; CR – correct rejection (true negative); FA – false alarm (false positive).

Fig. 5 ROC curve for all 1,768 phrases.

We also conducted χ2 test. Briefly, χ2 test is used with categorical data and can test for statistical independence of observed frequencies to what is expected. Here, observed frequencies are the counts of LLM “makes sense” and “nonsense” responses and the expected response frequencies are those provided by the human data (e.g., 977 nonsense and 761 meaningful). Table 3 shows that the LLM frequency of responses are significantly different from the human response frequencies, and supports SDT and ROC results.Table 3 A χ2 test results for observed (LLM) compared to expected (human) frequency of “makes sense” and “nonsense” responses. p < 0.05 indicates significantly different performance relative to humans.

LLM	χ2	p	
claude-3-Opus	9.4	p = 0.002	
gemini-1.0-Pro-001	87.0	p < 0.001	
gpt-3.5-turbo	1442.2	p < 0.001	
gpt-4-turbo	40.5	p < 0.001	

Table 2 and Figs. 4 and 5 show that GPT-3.5-turbo displays poor discrimination, and it judges most pairs as making sense (both high hit and false alarm rate). Gemini-1.0-Pro-001 and GPT-4-turbo display modest discrimination. They can correctly differentiate sensible and nonsense phrases while having a moderate chance of incorrectly judging nonsense phrases as making sense. Claude-3-Opus is substantially better than the other models and displays moderate-to-high discrimination abilities. Compared to GPT-4-turbo and Gemini-1.0-Pro-001, its ability to correctly detect nonsense and sensible phrases is further improved (higher correct rejection rate and lower false alarm rate). However, it is still significantly different from human performance.

Phrase-level frequency, semantic space, and meaningfulness ratings

The frequency with which humans encounter phrases is a vital part of ‘meaningfulness’, as phrases that are encountered more often will naturally be accepted as meaningful (although low-frequency phrases can also be meaningful). Similarly, meaningfulness ratings in LLMs could be expected to have a similar relationship with phrase-level frequency, due to their training corpora. We tested this by examining Spearman’s correlations between human/LLM TWT meaningfulness ratings and logarithmic Google bigram frequency (Log_Gfreq) for each phrase, as provided in the original Graves dataset. Both human and LLM ratings were strongly correlated with Log_Gfreq (except for GPT-3.5, which had a weak but still statistically significant correlation; Table 4).Table 4 Spearman’s correlations between TWT meaningfulness ratings and phrase-level Log_Gfreq.

Y: Log_Gfreq	
X: Rater	rs	p	
Humans	0.69	p < 0.001	
claude-3-Opus	0.66	p < 0.001	
gemini-1.0-Pro-001	0.63	p < 0.001	
gpt-3.5-turbo	0.13	p < 0.001	
gpt-4-turbo	0.67	p < 0.001	

Hence, both humans and LLMs appear to use bigram frequency to inform their meaningfulness judgments, and it is not clear where the discrepancy in ratings between humans and LLMs comes from. To identify possible reasons for the differences between human and LLM meaningfulness ratings, we tested the hypothesis that the LLMs are more likely to rate two words that are close together in semantic space as meaningful, regardless of whether the two words actually combine to form a coherent, meaningful phrase (e.g., the frog toad, the meat lamb). To test this, for each phrase, we calculated the semantic cosine similarity between the two words using 3 different semantic vector representations: GloVe, Word2Vec, and Taxonomic23–26. Using Spearman’s correlations, we investigated how meaningfulness judgments by humans/LLMs related to phrase-level semantic cosine similarity, controlling for Log_Gfreq (Table 5). Indeed, compared to humans, LLM meaningfulness ratings were more strongly correlated to semantic cosine similarity of the two words, suggesting that LLMs may have difficulty identifying when two highly semantically related words combine to create a low-meaningfulness phrase.Table 5 Spearman’s correlations between TWT meaningfulness ratings and phrase-level semantic cosine similarity.

X: Rater	rs	FDR p	
Y: GloVe	
Humans	0.03	p = 0.18	
claude-3-Opus	0.37	p < 0.001	
gemini-1.0-Pro-001	0.17	p < 0.001	
gpt-3.5-turbo	0.02	p = 0.35	
gpt-4-turbo	0.38	p < 0.001	
Y: Word2Vec	
Humans	0.10	p < 0.001	
claude-3-Opus	0.38	p < 0.001	
gemini-1.0-Pro-001	0.19	p < 0.001	
gpt-3.5-turbo	0.06	p = 0.01	
gpt-4-turbo	0.37	p < 0.001	
Y: Taxonomic	
Humans	− 0.01	p = 0.79	
claude-3-Opus	0.21	p < 0.001	
gemini-1.0-Pro-001	0.04	p = 0.18	
gpt-3.5-turbo	0.01	p = 0.84	
gpt-4-turbo	0.21	p < 0.001	
Significant correlation are indicated in bold.

Here, our task differs from common benchmark tasks that involve logic puzzles, arithmetic, games, or reasoning and inference given a sentence or short narrative27. In these problems, there is an objective “correct answer” that LLMs strive for, regardless of how well ordinary people perform on the problem. TWT, on the other hand, is based on subjective human judgments. The right answer for each word pair is what it is rated to be by a population of adults. For this task, the gold standard is the human rating, and there is no objectively true answer as such. Virtually any word pair can be said to “make sense” in some context or with sufficient mental gymnastics. A classic example of this is the poetry competition (http://archives.conlang.info/ga/farzhi/shiarweilwoen.html) in which many participants were able to write poems where the line “Colorless green ideas sleep furiously”—a line famously designed to not make sense—made sense. Given different instructions, humans can also come up with hypothetical, metaphorical, or poetic ‘meaningful’ interpretations for almost any two word phrase (e.g., the chair peacock is a species of peacock that likes to sit in chairs; the phone shoe is a shoe that can also serve as a phone). However, this is a different task than what is provided by the TWT, which constrains ‘meaningfulness’ via instructions and examples. TWT does not suggest that the low meaningfulness or nonsense word combinations cannot possibly make sense under any context. The claim is simply that some word combinations make much less sense on the surface or at face value, in the absence of any specific context. The task is to identify these combinations.

Limitations

Prompt design can have a profound impact on LLM performance. Here, we started by keeping the prompts as similar as possible to the original Graves et al.21 study, since the goal was to compare LLM performance to that of humans when given the same task. The prompt included multiple examples and was designed to be easy for humans to understand. LLMs are able to successfully complete similar tasks and instructions in other benchmarks. We performed a limited amount of prompt engineering, which included the addition of more examples (similar to the Few-Shot prompting with one example per rating value of the 4-point Likert scale), and a binary version of the task. However, it is possible that further modifications could tune LLM ratings to be closer to that of humans. Some popular prompting strategies (e.g., Chain-of-Thought and Tree-of-Thought), which are useful for improving the performance of complex tasks involving multiple reasoning steps, are not relevant to TWT since the problem does not contain any clear sub-goals or reasoning chains. A complete examination of the effects of prompt engineering on the TWT is beyond the scope of the current manuscript. However, future studies may experiment with instruction tuning to improve LLM performance on TWT by using the materials provided at https://github.com/NickRiccardi/two-word-test.

Due to the “black box” nature of LLMs tested here, it is unclear exactly why the LLMs provide meaningfulness judgments that are so different from humans. In the code provided, we have calculated similarity metrics for each phrase (cosine similarity of the two words in each phrase) based on some popular word embedding models such as word2vec. Future research could seek to find patterns using these metrics (or any number of other psycholinguistic properties) which may shed light on the instances in which LLMs are failing.

Finally, the Graves et al.21 dataset represents one of the largest sets of combinatorial semantic phrases that have human ratings, and the nature of its construction means that low- and high-sensibility phrases are closely matched to each other in important psycholinguistic variables. However, similar to many psycholinguistic studies, ratings were gathered in a relatively homogenous population (undergraduate students). Interesting directions for future research could include novel two-word phrases, metaphorical semantics, understanding combinations of more abstract concepts (e.g., love, freedom, etc.), or how ratings gathered from more diverse human populations may relate to LLM performance.

Conclusions and future work

We presented a new benchmark for testing language understanding in LLMs. The task, essentially trivial for humans, requires rating meaningfulness of two-word phrases. Four current LLMs fail on this task. While Claude-3-Opus performed better than GPT-4-turbo, Gemini-1.0-Pro-001, and GPT-3.5-turbo, its performance still fell well short of humans.

A binary version of the test, bTWT, was created to test whether the poor performance of LLMs was the result of a failure to deal with the numerical scale required for TWT. The bTWT revealed that GPT-3.5-turbo, Gemini and GPT-4-turbo fail to distinguish meaningfulness of phrases binarily, achieving poor-to-moderate discrimination. GPT-3.5-turbo was excessively liberal, tending to rate everything as “making sense”. Gemini-1.0-Pro-001 and GPT-4-turbo were able to detect the meaningfulness in some cases, while still having a higher chance of labeling nonsense phrases as “make sense”. Claude-3-Opus, however, takes a significant step forward on the bTWT.

Several investigations have begun to examine and reveal the limitations of LLMs. For example, Dziri et al.28 tested LLMs on three compositional tasks (multi-digit multiplication, logic grid puzzles, and dynamic programming). They found that LLMs solve compositional tasks by reducing multi-step compositional reasoning into linearized subgraph matching. They suggest that in multi-step reasoning problems, LLM performance will rapidly decay with increasing complexity. Failures have been demonstrated in other problems as well, such as those involving logical and common-sense reasoning29,30 as well as sequence tagging31.

The TWT differs from these cases in that it does not directly require inference or reasoning. The reaction time of making meaningfulness judgments in tasks similar to TWT is around 1 s for humans19,32, which suggests that it is a qualitatively different process than those used in typical benchmark tasks such as puzzle solving or logical and commonsense reasoning. A limitation in breaking down a complex chain of reasoning into smaller problems should not affect performance on the TWT. Understanding these phrases requires understanding the constituent concepts, and then using world knowledge to determine whether the combination makes sense in some manner. A “mountain stream” is a stream located on a mountain, but a “stream mountain” is not a thing at all. An “army knife” is not necessarily a knife located in the army but a type of knife useful in certain situations. TWT may exploit the fact that the text corpora that LLMs are trained on, no matter how large, almost entirely contain sensible text. However, this is the case for humans as well. Almost all text that people are exposed to is also sensible, but if the task requires, they are easily able to determine that certain word combinations don’t make much sense. Current LLMs may lack the depth of real-world knowledge that is required for this task.

Many of the limitations of LLMs identified previously can be associated with a lack of “executive control” that presents difficulties in complex symbolic or rule-based reasoning. Because of this, many have proposed combining deep neural networks with symbolic reasoning systems that can exert executive control when required (e.g., in three-digit multiplication). The weakness identified by TWT is qualitatively distinct, in that it is not directly related to the ability for executive control or systematic application of rules.

The importance of LLMs largely centers around the idea that they can understand language. The notion of understanding and meaningfulness implies that not everything “makes sense.” As an analogy, the notion of grammaticality is based on the fact that not every sentence is grammatical–certain sentences and phrases fall outside the grammaticality boundary, which linguists commonly attempt to delineate. Similarly, the notion of semantics also implies that there are certain things that do not make sense or are incoherent. In this context, there are two broad classes of applications of LLMs. In the first class, there is a clearly defined correct answer where one wants the LLM to produce that answer regardless of what humans do for such problems, as mentioned above. Examples are logical reasoning problems or generating code. The second class of applications is where one would want LLMs to behave like humans. In these cases, the goal is for LLM’s sense of understanding or meaning to be similar to that of humans. Examples of such cases are an LLM avatar attending a meeting in place of an employee, providing customer support in relation to some products or services, or automatically replying to emails. Here, one would desire for the LLM to behave like humans. This would include detecting cases where messages don’t make sense, and where appropriate actions would be asking for clarifications, correcting a mistake, or denying a request, instead of creatively interpreting every message as sensible. Nonsensical messages can arise in real-world scenarios due to many reasons, including mistakes, misunderstandings, disorders, or malicious attacks. We believe that TWT is most relevant for these second classes of problems. It allows us to test whether LLMs have a boundary that is similar to that of humans in relation to meaningfulness.

These results also urge for caution in the attribution of AGI or similar abilities to LLMs, based only on testing on tasks that are difficult for humans. Seemingly simple language tasks can be difficult for LLMs even when they perform well on complex tasks, which is reminiscent of Moravec's paradox. The mounting understanding of the impressive abilities as well as limitations of LLMs will be essential in improving these models, and in identifying appropriate use cases.

Supplementary Information

Supplementary Information.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-024-72528-3.

Acknowledgements

This work was supported by NIH/NIDCD R01 DC017162 (RHD).

Author contributions

N.R. and R.H.D. conceptualized the research. N.R. and X.Y. conducted the experiments, analyzed the data, and generated the visualizations. N.R. wrote the first draft. All authors revised and reviewed the manuscript.

Funding

National Institutes of Health, NIDCD, R01DC017162.

Data availability

The data generated by LLMs that was used in this study are available at: https://github.com/NickRiccardi/two-word-test.

Competing interests

The authors declare no competing interests.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Bommasani, R. et al. On the opportunities and risks of foundation models. (2021) 10.48550/ARXIV.2108.07258.
2. Choi JH Hickman KE Monahan A Schwarcz DB ChatGPT goes to law school SSRN J. 2023 10.2139/ssrn.4335905
Choi, J. H., Hickman, K. E., Monahan, A. & Schwarcz, D. B. ChatGPT goes to law school. SSRN J.10.2139/ssrn.4335905 (2023).10.2139/ssrn.4335905
3. Terwiesch, C. Would Chat GPT get a Wharton MBA? A prediction based on its performance in the operations management course. (2023).
4. Kung TH Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models PLOS Digit Health 2023 2 e0000198 10.1371/journal.pdig.0000198 36812645
Kung, T. H. et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLOS Digit Health 2, e0000198 (2023).36812645 10.1371/journal.pdig.0000198
5. Brown, T. B. et al. Language models are few-shot learners. (2020) 10.48550/ARXIV.2005.14165.
6. Chowdhery, A. et al. PaLM: Scaling language modeling with pathways. (2022) 10.48550/ARXIV.2204.02311.
7. Bubeck, S. et al. Sparks of artificial general intelligence: Early experiments with GPT-4. (2023) 10.48550/ARXIV.2303.12712.
8. Michael, J. et al. What Do NLP Researchers Believe? Results of the NLP Community Metasurvey. (2022) 10.48550/ARXIV.2208.12852.
9. Mitchell M Krakauer DC The debate over understanding in AI’s large language models Proc. Natl. Acad. Sci. U.S.A. 2023 120 e2215907120 10.1073/pnas.2215907120 36943882
Mitchell, M. & Krakauer, D. C. The debate over understanding in AI’s large language models. Proc. Natl. Acad. Sci. U.S.A. 120, e2215907120 (2023).36943882 10.1073/pnas.2215907120
10. Lee, K., Firat, O., Agarwal, A., Fannjiang, C. & Sussillo, D. Hallucinations in neural machine translation. in (2018).
11. Raunak, V., Menezes, A. & Junczys-Dowmunt, M. The Curious Case of Hallucinations in Neural Machine Translation. in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies 1172–1183 (Association for Computational Linguistics, Online, 2021). 10.18653/v1/2021.naacl-main.92.
12. Mahowald, K. et al. Dissociating language and thought in large language models. (2023) 10.48550/ARXIV.2301.06627.
13. Choudhury, S. R., Rogers, A. & Augenstein, I. Machine Reading, Fast and Slow: When Do Models ‘Understand’ Language? (2022) 10.48550/ARXIV.2209.07430.
14. Gardner, M. et al. Competency problems: On finding and removing artifacts in language data. (2021) 10.48550/ARXIV.2104.08646.
15. Linzen, T. How can we accelerate progress towards human-like linguistic generalization? (2020) 10.48550/ARXIV.2005.00955.
16. Browning, J. & Lecun, Y. AI and the limits of language. Noema (2022).
17. Gagné CL Relation and lexical priming during the interpretation of noun–noun combinations J. Exp. Psychol. Learn. Mem. Cognit. 2001 27 236 254 10.1037/0278-7393.27.1.236 11204100
Gagné, C. L. Relation and lexical priming during the interpretation of noun–noun combinations. J. Exp. Psychol. Learn. Mem. Cognit. 27, 236–254 (2001).11204100 10.1037/0278-7393.27.1.236
18. Gagné CL Spalding TL Constituent integration during the processing of compound words: Does it involve the use of relational structures? J. Mem. Lang. 2009 60 20 35 10.1016/j.jml.2008.07.003
Gagné, C. L. & Spalding, T. L. Constituent integration during the processing of compound words: Does it involve the use of relational structures?. J. Mem. Lang. 60, 20–35 (2009).10.1016/j.jml.2008.07.003
19. Graves WW Binder JR Desai RH Conant LL Seidenberg MS Neural correlates of implicit and explicit combinatorial semantic processing NeuroImage 2010 53 638 646 10.1016/j.neuroimage.2010.06.055 20600969
Graves, W. W., Binder, J. R., Desai, R. H., Conant, L. L. & Seidenberg, M. S. Neural correlates of implicit and explicit combinatorial semantic processing. NeuroImage 53, 638–646 (2010).20600969 10.1016/j.neuroimage.2010.06.055
20. Pylkkänen L The neural basis of combinatory syntax and semantics Science 2019 366 62 66 10.1126/science.aax0050 31604303
Pylkkänen, L. The neural basis of combinatory syntax and semantics. Science 366, 62–66 (2019).31604303 10.1126/science.aax0050
21. Graves WW Binder JR Seidenberg MS Noun–noun combination: Meaningfulness ratings and lexical statistics for 2,160 word pairs Behav Res 2013 45 463 469 10.3758/s13428-012-0256-3
Graves, W. W., Binder, J. R. & Seidenberg, M. S. Noun–noun combination: Meaningfulness ratings and lexical statistics for 2,160 word pairs. Behav Res 45, 463–469 (2013).10.3758/s13428-012-0256-3
22. Crawford JR Howell DC Comparing an individual’s test score against norms derived from small samples Clin. Neuropsychol. 1998 12 482 486 10.1076/clin.12.4.482.7241
Crawford, J. R. & Howell, D. C. Comparing an individual’s test score against norms derived from small samples. Clin. Neuropsychol. 12, 482–486 (1998).10.1076/clin.12.4.482.7241
23. Mikolov, T., Chen, K., Corrado, G. & Dean, J. Efficient Estimation of Word Representations in Vector Space. Preprint at 10.48550/ARXIV.1301.3781 (2013).
24. Pennington, J., Socher, R. & Manning, C. Glove: Global Vectors for Word Representation. in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) 1532–1543 (Association for Computational Linguistics, Doha, Qatar, 2014). 10.3115/v1/D14-1162.
25. Roller, S. & Erk, K. Relations such as Hypernymy: Identifying and Exploiting Hearst Patterns in Distributional Vectors for Lexical Entailment. in Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing 2163–2172 (Association for Computational Linguistics, Austin, Texas, 2016). 10.18653/v1/D16-1234.
26. Gao C Shinkareva SV Desai RH SCOPE: The South Carolina psycholinguistic metabase Behav Res 2022 55 2853 2884 10.3758/s13428-022-01934-0
Gao, C., Shinkareva, S. V. & Desai, R. H. SCOPE: The South Carolina psycholinguistic metabase. Behav Res 55, 2853–2884 (2022).10.3758/s13428-022-01934-0
27. Srivastava, A. et al. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models. (2022) 10.48550/ARXIV.2206.04615.
28. Dziri, N. et al. Faith and Fate: Limits of Transformers on Compositionality. (2023) 10.48550/ARXIV.2305.18654.
29. Bian, N. et al. ChatGPT is a Knowledgeable but Inexperienced Solver: An Investigation of Commonsense Problem in Large Language Models. (2023) 10.48550/ARXIV.2303.16421.
30. Koralus, P. & Wang-Maścianica, V. Humans in Humans Out: On GPT Converging Toward Common Sense in both Success and Failure. (2023) 10.48550/ARXIV.2303.17276.
31. Qin, C. et al. Is ChatGPT a General-Purpose Natural Language Processing Task Solver? in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing 1339–1384 (Association for Computational Linguistics, Singapore, 2023). 10.18653/v1/2023.emnlp-main.85.
32. Parrish A Pylkkänen L Conceptual combination in the LATL with and without syntactic composition Neurobiol. Lang. 2022 3 46 66 10.1162/nol_a_00048
Parrish, A. & Pylkkänen, L. Conceptual combination in the LATL with and without syntactic composition. Neurobiol. Lang. 3, 46–66 (2022).10.1162/nol_a_00048
