
==== Front
PLoS One
PLoS One
plos
PLOS ONE
1932-6203
Public Library of Science San Francisco, CA USA

PONE-D-24-26240
10.1371/journal.pone.0310730
Research Article
Research and Analysis Methods
Imaging Techniques
Physical Sciences
Physics
Thermodynamics
Entropy
Computer and Information Sciences
Artificial Intelligence
Machine Learning
Computer and Information Sciences
Software Engineering
Preprocessing
Engineering and Technology
Software Engineering
Preprocessing
Computer and Information Sciences
Information Technology
Natural Language Processing
Biology and Life Sciences
Neuroscience
Cognitive Science
Cognitive Psychology
Language
Biology and Life Sciences
Psychology
Cognitive Psychology
Language
Social Sciences
Psychology
Cognitive Psychology
Language
Biology and Life Sciences
Organisms
Eukaryota
Animals
Vertebrates
Amniotes
Mammals
Dogs
Biology and Life Sciences
Zoology
Animals
Vertebrates
Amniotes
Mammals
Dogs
Physical Sciences
Mathematics
Statistics
Similarity Measures
Cosine Similarity
Distilling knowledge from multiple foundation models for zero-shot image classification
Distilling knowledge from multiple foundation models for zero-shot image classification
https://orcid.org/0009-0006-5193-2379
Yin Siqi Methodology 1 *
Jiang Lifan Writing – review & editing 2
1 School of Computer Science and Technology, Shandong University of Science and Technology, Qingdao, Shandong, China
2 School of Computer Science and Technology, Shandong University of Science and Technology, Qingdao, Shandong, China
Chaudhary Sushank Editor
Guangdong University of Petrochemical Technology, CHINA
Competing Interests: The authors have declared that no competing interests exist.

* E-mail: siqiyin@sdust.edu.cn, siqiyin123@gmail.com
2024
20 9 2024
19 9 e031073027 6 2024
5 9 2024
© 2024 Yin, Jiang
2024
Yin, Jiang
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Zero-shot image classification enables the recognition of new categories without requiring additional training data, thereby enhancing the model’s generalization capability when specific training are unavailable. This paper introduces a zero-shot image classification framework to recognize new categories that are unseen during training by distilling knowledge from foundation models. Specifically, we first employ ChatGPT and DALL-E to synthesize reference images of unseen categories from text prompts. Then, the test image is aligned with text and reference images using CLIP and DINO to calculate the logits. Finally, the predicted logits are aggregated according to their confidence to produce the final prediction. Experiments are conducted on multiple datasets, including MNIST, SVHN, CIFAR-10, CIFAR-100, and TinyImageNet. The results demonstrate that our method can significantly improve classification accuracy compared to previous approaches, achieving AUROC scores of over 96% across all test datasets. Our code is available at https://github.com/1134112149/MICW-ZIC.

The author(s) received no specific funding for this work. Data AvailabilityAll files are available from these database: https://doi.org/10.6084/m9.figshare.26815960.v1 https://doi.org/10.6084/m9.figshare.26815945.v2.
Data Availability

All files are available from these database: https://doi.org/10.6084/m9.figshare.26815960.v1 https://doi.org/10.6084/m9.figshare.26815945.v2.
==== Body
pmc1 Introduction

In recent years, deep learning technologies have made significant progresses, particularly in the field of image classification, with numerous successful models being developed. For instance, ResNet [1] introduced a residual learning mechanism to address the difficulties in training deep networks, substantially increasing classification accuracy. Meanwhile, MobileNet [2] achieved efficient computation on mobile and embedded devices with its lightweight deep separable convolution design. Despite the significant success, these methods are built upon an assumption that the categories of the test images have been seen in the training set, which imposes restrictions in their practical applications. In real-world scenarios, we often encounter a large number of new categories that may not be included in the original training set.

To overcome the limitations of the aforementioned methods when dealing with unseen categories, zero-shot learning has garnered increasing attention in recent years, which aims to enable models to recognize categories that they have never seen during the training phase. Several prominent zero-shot models have been developed from various aspects. For example, an early zero-shot method [3] improved the classification accuracy of unseen bird species by learning the mapping between visual features and textual descriptions. Due to the relatively small number of images for training models, these methods still face challenges in achieving satisfactory zero-shot generalization capabilities.

In this paper, we propose a novel framework to leverage the alignment of images with text as well as the alignment among images themselves to further improve the zero-shot accuracy using multiple foundation models like ChatGPT, CLIP and DINO. Specifically, we first employ ChatGPT and DALL-E to generate images of unseen categories from text prompts. Then, the test image is aligned with text and reference images using CLIP and DINO to calculate the logits. Finally, the predicted logits are aggregated according to their confidence to produce the final prediction. To evaluate the effectiveness of our method, we conduct experiments on multiple datasets including CIFAR-10, CIFAR-100, and TinyImageNet. In comparison to previous zero-shot methods, our model produces significant gains in classification accuracy, with improvements of 11.97%, 26.06%, and 12.6% on CIFAR-10; 24.48%, 42.48%, and 5.42% on CIFAR-100; and 32.32%, 41.96%, and 1.2% on TinyImageNet, respectively. Furthermore, our model produces over 96% in terms of AUROC across all test datasets, particularly exceeding 99% on the CIFAR-10 dataset. These results clearly demonstrate the superiority of our method in zero-shot image classification.

2 Related work

2.1 Zero-shot classification

Zero-shot classification (ZSC) is a machine learning paradigm aimed at enabling models to recognize categories they have not seen during the training phase. Unlike traditional supervised learning methods, ZSC does not rely on direct experience with every class but rather achieves classification through understanding shared knowledge or attributes among categories [4–8]. This capability is crucial for dealing with data scarcity or category diversity issues, especially in fields such as Natural Language Processing (NLP) and Computer Vision (CV). In recent years, significant progress has been made in the field of ZSC. In particular, pre-trained language models such as GPT-3 [9] and BERT [10] show great potential in ZSC tasks. These models are trained on massive text datasets to learn rich linguistic features and world knowledge, which enables them to infer unknown categories without explicit examples. Additionally, by operating on graph-structured data, GNNs is able to capture complex relationships and attributes between categories, offering new avenues for accurate classification of unseen classes [11, 12]. To further improve zero-shot performance, Deep Calibration Network (DCN) [13] is developed to reduce the uncertainty of models when dealing with unseen categories. In addition, meta-learning has been studied to facilitate models to achieve higher accuracy on unseen images [14]. Moreover, a framework based on Generative Adversarial Networks (GANs) is proposed in [15], which can generate features representative of unseen categories to achieve further gains. To address the data imbalance issue in zero-shot classification, Balanced Meta-Softmax [16] is developed to achieve a better balance between different categories. Recent advancements such as [17], which refines the alignment between semantic and visual spaces, and [18], which improves disentangling cluster-based semantic representations, have further enhanced ZSC performance.

With the rapid advancements of recent foundation models, numerous attempts have been made to study these models from the perspective of zero-shot classification. CLIP [19] has demonstrated its superior performance across multiple computer vision benchmarks, particularly showing promising accuracy in zero-shot classification tasks. The success of CLIP highlights the potential of learning visual tasks with language supervision. Motivated by this, several attempts [20–24] have been made to adopt CLIP to classify images of unseen categories and produce promising results. For example, models like Calip [24] and ZegCLIP [23] leverage cross-modal joint learning to achieve more accurate classification of unseen categories by better aligning visual and textual representations. However, while these methods improve classification accuracy, they still face challenges in fully capturing the nuances of diverse and complex unseen categories, particularly in scenarios with limited or noisy data. Other methods [25–27] employ DINO to leverage the superior generalization capability of its features to boost downstream tasks like image classification. Recently, generative foundation models like DALL-E are also investigated for the zero-shot classification task. Specifically, Zhang et al. [28] employ generative models to produce additional images to augment the training dataset to achieve accuracy gains in zero-shot classification.

Despite superior performance of the aforementioned methods, the rich knowledge in foundation models are not fully exploited. To remedy this, in this paper, we propose to leverage foundation models from two perspectives to make full use of their knowledge. First, we employ ChatGPT and DALL-E to synthesize reference images that precisely describe unseen categories and classification boundaries, obtaining reference samples for new categories. Then, CLIP and DINO are employed to align the test image with reference images in the representation space to conduct image-image alignment. Meanwhile, CLIP is also adopted to associate the test image with text prompts for text-image alignment. Through these two kinds of alignment, the knowledge of foundation models can be better exploited to produce higher accuracy. Note that, different from [28] that uses generative models to produces additional samples for data augmentation, we employ generative models to synthesize reference images that are usually confused in appearance. As a result, our method achieves significant performance gains as compared to previous approaches, producing an increase of 0.45%, 0.24%, and 0.23% in classification accuracy on the CIFAR10, CIFAR100, and TinyImageNet datasets, respectively.

3 Method

The overview of our method is summarized in Fig 1. Since GPT [29, 30] demonstrates its excellent capabilities in text generation and question answering, it is employed for text understanding. Inspired by the remarkable performance of DALL-E [31], it is adopted for image generation in our framework. To achieve text-image alignment, CLIP [19] is used due to its powerful capability. To extract distinct representation from the image, DINO [25] is selected due to its superior performance. Assume we have a dataset (D) containing m seen classes and n unseen classes. When conducting an image classification task on the dataset D, we first generate reference images for each category. Then, the test image x is aligned with class-definition text using CLIP to calculate the logits. Meanwhile, the test image is also aligned with reference images using CLIP and DINO to obtain the logits. Finally, the resultant logits are aggregated according to their confidence to obtain the final prediction.

10.1371/journal.pone.0310730.g001 Fig 1 Overview of our method.

3.1 Reference images generation

To synthesize images of unseen categories, ChatGPT and DALL-E are employed to leverage their rich knowledge.

Note that, several categories are easily confused in appearance since they may be similar in shape or features. Thus, we propose to generate reference images that are located near the boundaries of ambiguous categories. Specifically, we first conduct feature analysis to obtain common features of similar categories. Then, we synthesize images based on the resultant common features.

(1) Feature analysis

First, ChatGPT is adopted to analyze the appearance of m + n categories in the dataset, as shown in the left column of Fig 2. Then, ChatGPT is used to group categories that are similar in appearance. Specifically, we ask ChatGPT to identify those similar categories that are easily confused. Next, for each pair of resultant groups, ChatGPT is further adopted to analyze their common appearance features. Finally, the common features are obtained and denoted as F.

10.1371/journal.pone.0310730.g002 Fig 2 Examples of utilizing ChatGPT to generate common features of similar categories.

(2) Image synthesis

After feature analysis, DALL-E is employed to generate images of each class that possess the common appearance features F. For categories that do not resemble other categories, we directly utilize DALL-E to generate images belonging to it without using detailed appearance descriptions. To acquire more comprehensive reference information, we generate multiple images for each category. All these images will be employed in the subsequent steps for image-image alignment.

3.2 Alignment-driven image classification

To achieve accurate classification of an unseen test image, we align the image with text prompt and reference images using foundation models to distill their knowledge.

3.2.1 Text-image alignment

CLIP is capable of encoding image and text into a shared representational space. With this property, the logits of the test image x and the unseen category tn can be obtained by encoding x and tn using CLIP and calculating the cosine similarity S of the resultant representations: Si-t,clipn=cos(Fi(x),Ft(tn)), (1)

where Fi(⋅) and Ft(⋅) represent the image encoder and text encoder of CLIP.

3.2.2 Image-image alignment

In addition to text-image alignment, we also conduct image-image alignment to leverage powerful vision foundation models. Specifically, the test image and the reference images are first encoded using CLIP and DINO to obtain their representations. Then, the cosine similarities between the resultant representations are calculated: Si-i,clipn=1M∑i=1Mcos(Fi(x),Fi(rn)), (2)

Si-i,dinon=1M∑i=1Mcos(D(x),D(rn)), (3)

where D(⋅) refers to the image encoder of DINO, rn denotes the nth reference image.

3.3 Logits aggregation

After text-image and image-image alignment, a confidence-based strategy is employed to aggregate the resultant logits. That is, the logits with higher confidence are assigned with larger weights to produce the final results. Next, the final output is obtained as: Sfuse=Wi-t,clip·Si-t,clip+Wi-i,clip·Si-i,clip+Wi-i,dino·Si-i,dino. (4)

The confidence coefficient W is calculated using the following three options, which will be studied through ablation experiments in Sec. 4.3.

(1) Using the maximum value of each S as the confidence level: W=max{Sn}n=1N. (5)

(2) Using the inverse of the entropy as the confidence value: W=1H+10-6,H=-∑nNSnlog(Sn). (6)

where 10−6 is employed to avoid the denominator to be 0.

(3) Using the negative exponential of the entropy as the confidence value: W=exp(-H),H=-∑nNSnlog(Sn). (7)

4 Experiments

4.1 Datasets and evaluation metrics

To evaluate the effectiveness of our proposed method, we conduct experiments on five widely-used datasets following previous protocol [32–34], including MNIST, SVHN, CIFAR10, CIFAR100, and TinyImageNet. For MNIST and SVHN and CIFAR10, we all select 6 categories as the closed set and the remaining 4 categories as the open set. For the CIFAR+N experiments, 4 classes are randomly sampled as the closed-set classes and N non-overlapping classes are randomly sampled from CIFAR100 as the open-set classes for evaluation (N = 10 for CIFAR-10 and N = 50 for CIFAR+50). For TinyImageNet, we select 20 categories as the closed set, with the other 180 categories as the open set. We employ Accuracy (AC) and Area Under the Receiver Operating Characteristic curve (AUROC). The AUROC metric is used to evaluate the model’s ability to differentiate between closed-set and open-set samples, that is, the accuracy of the model in identifying novel category samples.

4.2 Main results

We compare our method with nine state-of-the-art methods, including RPL [35], OpenHybrid [34], PMAL [36], ZOC [37], MLS [38], DIAS [39], Class-iinclusion [40], ODL [41], and CSSR [42]. Quantitative results are presented in Table 1. Compared with previous open-set classification methods, our method produces significant accuracy gains, especially on CIFAR-10 and TinyImageNet. For example, our method outperforms Class-inclusion by over 5% and 3% on CIFAR-10 and TinyImageNet, respectively. The higher accuracy of our method clearly demonstrate its effectiveness and superiority.

10.1371/journal.pone.0310730.t001 Table 1 AUROC results achieved by different methods.

Method	MNIST	SVHN	CIFAR10	CIFAR+10	CIFAR+50	TinyImageNet	
Methods that involve a training process	
RPL [35]	91.7	93.1	90.1	97.6	96.8	80.9	
OpenHybrid [34]	99.5	94.7	95.0	96.2	95.5	79.3	
PMAL [36]	99.7	97.0	95.1	97.8	96.9	83.1	
ZOC [37]	/	/	93.0	97.8	97.6	84.6	
MLS [38]	99.3	97.1	93.6	97.9	96.5	83.0	
DIAS [39]	99.2	94.3	85.0	92.0	91.6	73.1	
Class-inclusion [40]	/	95.6	94.8	95.7	95.7	78.5	
ODL [41]	99.5	94.3	85.7	89.1	88.3	76.4	
ODL+ [41]	99.6	95.4	88.5	91.1	90.6	74.6	
CSSR [42]	/	97.9	91.3	96.3	96.2	82.3	
RCSSR [42]	/	97.8	91.5	96.0	96.3	81.9	
Methods that involve no extra training process	
Ours	99.7	97.8	99.8	98.4	97.8	96.4	

Table 2 shows the performance of our method compared to other methods across CIFAR10, CIFAR100, and TinyImageNet datasets. Our method achieves the best results in all metrics, including Top1, Top3, Top5 accuracy, and AUROC. Specifically, on CIFAR10, our method reaches 91.96% in terms of Top1 accuracy and 99.78% on AUROC. Similarly, it delivers the highest performance on CIFAR100 and TinyImageNet datasets.

10.1371/journal.pone.0310730.t002 Table 2 Performance comparison of single methods and our three-method fusion across CIFAR10, CIFAR100, and TinyImageNet datasets.

Method	CIFAR10	CIFAR100	TinyImageNet	
Top1	Top3	Top5	AUROC	Top1	Top3	Top5	AUROC	Top1	Top3	Top5	AUROC	
CLIP_text	79.99%	93.22%	97.41%	97.26%	47.69%	66.14%	72.35%	91.41%	41.20%	61.16%	69.24%	85.95%	
CLIP_image	65.90%	86.20%	86.20%	98.12%	30.43%	48.55%	56.67%	85.73%	31.56%	48.17%	55.57%	82.68%	
DINO	79.36%	90.48%	94.51%	97.97%	66.75%	81.77%	86.48%	94.28%	72.32%	72.32%	86.79%	96.07%	
Ours	91.96%	98.14%	99.30%	99.78%	72.17%	86.55%	90.36%	96.03%	73.52%	84.95%	88.55%	96.48%	

4.3 Ablation study

4.3.1 Effectiveness of different components

We first conduct experiments to study the effectiveness for different components of our method, including image-text alignment using CLIP (M1), image-image alignment using CLIP (M2)m and image-image alignment using DINO (M3). Specifically, we compare the performance of our method with different combinations of components. As shown in Tables 3 and 4, all the components in our methods contribute to higher accuracy. Without M1/M2/M3, the network variant suffers notable accuracy drop on most metrics. Moreover, M1 and M3 contribute more significantly than M2. These results clearly demonstrate the effectiveness of each component in our method.

10.1371/journal.pone.0310730.t003 Table 3 Top1, Top3, Top5 results achieved by our method with different settings.

Method	CIFAR10	CIFAR100	TinyImageNet	
Top1	Top3	Top5	Top1	Top3	Top5	Top1	Top3	Top5	
M1+M2	83.83%	94.85%	98.22%	48.00%	66.41%	73.27%	43.52%	63.97%	71.79%	
M1+M3	92.51%	98.16%	99.35%	71.93%	86.59%	90.50%	73.29%	84.73%	88.41%	
M2+M3	79.99%	93.22%	97.41%	64.19%	79.29%	84.63%	72.74%	84.33%	87.30%	
Ours (M1+M2+M3)	92.96%	98.32%	99.33%	72.17%	86.55%	90.36%	73.52%	84.95%	88.55%	

10.1371/journal.pone.0310730.t004 Table 4 AUROC results achieved by our method with different settings.

Method	CIFAR10	CIFAR100	TinyImageNet	
M1+M2	97.94%	92.99%	87.03%	
M1+M3	99.75%	95.87%	96.54%	
M2+M3	97.26%	94.70%	96.09%	
Ours (M1+M2+M3)	99.78%	96.03%	96.48%	

4.3.2 Effects of different aggregation strategies

We further conduct experiments to study different aggregation strategies. Specifically, we first use predefined weights to aggregate the logits produced by different methods. Then, we study three different weighting strategies, which employ max similarity, inverse entropy, and negative exponential of entropy to fuse logits. As shown in Tables 5 and 6, the method with inverse entropy produces the best performance. Consequently, this strategy is used as the default setting in our experiments.

10.1371/journal.pone.0310730.t005 Table 5 Top1, Top3, Top5 accuracy achieved by our method with different weight settings on CIFAR10.

Weighting Method	Top1	Top3	Top5	
1:1:1	92.36%	98.22%	99.31%	
3:3:4	92.29%	98.21%	99.26%	
Max Similarity	92.75%	98.26%	99.34%	
Inverse Entropy	92.96%	98.32%	99.33%	
Negative Exponential of Entropy	92.90%	98.31%	99.36%	

10.1371/journal.pone.0310730.t006 Table 6 AUROC results achieved by our method with different weight setting methods on CIFAR10.

Weighting Method	AUROC	
1:1:1	99.55%	
3:3:4	99.60%	
Max Similarity	99.68%	
Inverse Entropy	99.78%	
Negative Exponential of Entropy	99.73%	

4.3.3 Quantity and quality of generated reference images

Intuitively, the quantity and quality of generated reference images are critical to the final performance. Consequently, we conducted experiments on generating reference images. To increase the diversity of generated images, the object with various realistic poses were generated to depict activities like cats eating, sleeping, and standing. As illustrated in Fig 3, the accuracy (ACC-TOP1) progressively increased from 67% to 73% with more images being generated. This demonstrates large quantity is beneficial to the final accuracy. In addition, we constructed a network variant with only one reference image per category and compared its accuracy with our baseline model with multiple reference images. As shown in Table 7, the model with multiple reference images significantly outperforms the one with a single reference image. This is because additional reference images provide more comprehensive information to describe the properties of each category, ultimately leading to higher accuracy.

10.1371/journal.pone.0310730.g003 Fig 3 Ablation of the reference images.

10.1371/journal.pone.0310730.t007 Table 7 Comparison of Top1 and AUROC results when using one and multiple reference images.

The left side of each ‘-’ represents the Top1 result, and the right side of each ‘-’ represents the AUROC result.

Number of Reference Images	CIFAR10	CIFAR100	TinyImageNet	
One	88.71%-99.05%	67.78%-96.025%	67.77%-96.00%	
Multiple	92.96%-99.78%	72.17%-96.026%	73.52%-96.48%	

On the other hand, as shown in Fig 3, images of low quality may cause confusions between the true category, cat, and other categories like lesser panda, deer, fox, and dog. As a result, ACC-TOP1 is decreased to 64%. This demonstrates that the quality of generated images is also crucial. It can be observed from Fig 4 that the generated hard samples facilitate the model to better distinguish similar categories, thereby achieving higher accuracy.

10.1371/journal.pone.0310730.g004 Fig 4 Visualization of synthesized reference images.

4.3.4 Visualization results

We first visualize the synthetic images in Fig 4. It can be observed that our method is able to synthesize images for similar categories with similar appearance, which facilitate our method to produce higher accuracy.

We then conduct a detailed analysis for the features of two similar groups of categories [cats, dogs] and [cars, trucks] within the CIFAR-10 dataset. Concretely, for each category, we visualize the features of all test images and the corresponding synthetic images. We employ the t-SNE technique to visualize the representations in Fig 5. As shown in Fig 5, we observe that feature points of synthetic reference images are located at the boundaries between categories, which facilitate the model to distinguish similar categories more clearly.

10.1371/journal.pone.0310730.g005 Fig 5 Visualization of the representations for test images and reference images using t-SNE.

4.3.5 Analysis of computational complexity and scalability

As shown in Table 8, we present a comprehensive breakdown of the inference time for each stage of our method. Notably, the runtime for steps (1) and (2) is excluded since these steps are conducted using webpage APIs. It is clear that the inference time for steps (3) to (7) is relatively small. For instance, image preprocessing takes approximately 2.1 × 10−4 seconds, while feature extraction and classification require 1.3 × 10−2 seconds and 2.6 × 10−2 seconds, respectively. Overall, processing the entire CIFAR-10 dataset requires approximately 232 seconds, which shows the efficiency of our approach. For larger datasets such as CIFAR-100 and TinyImageNet, the total processing times are approximately 430 seconds and 1970 seconds, respectively, remaining within a reasonable computational range.

10.1371/journal.pone.0310730.t008 Table 8 Timeline for zero-shot image classification task on CIFAR-10 dataset.

Step	Description	Time (seconds)	
(1)	GPT: generate unknown labels based on known labels	/	
(2)	GPT: combine similar categories and feature analysis	/	
(3)	DALL-E: generate reference images	15 – 60	
(4)	Test-images preprocessing	2.1 × 10−4	
(5)	CLIP and DINO: Feature extraction	1.3 × 10−2	
(6)	Classification	2.6 × 10−2	
(7)	Post-processing	1.2 × 10−4	

5 Conclusion

In this paper, we introduce a framework for zero-shot classification by exploiting knowledge from multiple foundation models. We first employ ChatGPT and DALL-E to synthesize reference images of unseen categories. Then, text-image alignment and image-image alignment are conducted to produce the logits of the test image. Finally, the resultant logits are aggregated to generate the overall prediction. Extensive experimental results on CIFAR-10, CIFAR-100, and TinyImageNet validate the effectiveness of our approach, showcasing its exceptional performance in both open and closed set classification tasks.

Despite promising performance, our method still has some limitations. Specifically, the process of generating synthetic images is relatively computationally complex, especially when dealing with large-scale datasets or real-time applications, which could lead to significant computational overhead. In the future, we will focus on improving the computational efficiency of the image generation process to better fit real-world applications.
==== Refs
References

1 He, K. and Zhang, X. and Ren, S. and Sun, J., Deep residual learning for image recognition, Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778, 2016.
2 Howard, A. G. and Zhu, M. and Chen, B. and Kalenichenko, D. and Wang, W. and Weyand, T. et al., MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications, arXiv preprint arXiv:1704.04861, 2017.
3 Paz-Argaman, T. and Atzmon, Y. and Chechik, G. and Tsarfaty, R., Zest: Zero-Shot Learning from Text Descriptions Using Textual Similarity and Visual Summarization, arXiv preprint arXiv:2010.03276, 2020.
4 Naeem, M. F., Xian, Y. Q., Tombari, F., and Akata, Z. Learning graph embeddings for compositional zero-shot learning, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 953–962, 2021.
5 Wang W. , Zheng V. W. , Yu H. , and Miao C. Y . A survey of zero-shot learning: Settings, methods and applications, ACM Transactions on Intelligent Systems and Technology (TIST), vol. 10 , no. 2 , pp. 1–37, 2019. doi: 10.1145/3293318
6 McCartney B. , Martinez-del-Rincon J. , Devereux B. , and Murphy B . A zero-shot learning approach to the development of brain-computer interfaces for image retrieval, PLoS One, vol. 14 , no. 9 , pp. e0214342, 2019. doi: 10.1371/journal.pone.0214342 31525201
7 Auzepy A. , Tönjes E. , Lenz D. , and Funk C . Evaluating TCFD reporting—A new application of zero-shot analysis to climate-related financial disclosures, PLoS One, vol. 18 , no. 11 , pp. e0288052, 2023. doi: 10.1371/journal.pone.0288052 37917605
8 Lu, J., Li, J., Yan, Z., and Zhang, C. S. Zero-shot learning by generating pseudo feature representations, arXiv preprint arXiv:1703.06389, 2017.
9 Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., et al. Language models are few-shot learners, Advances in Neural Information Processing Systems, vol. 33, pp. 1877–1901, 2020.
10 Devlin, J., Chang, M. W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding, arXiv preprint arXiv:1810.04805, 2018.
11 Ou G. , Yu G. , Domeniconi C. , Lu X. , and Zhang X . Multi-label zero-shot learning with graph convolutional networks, Neural Networks, vol. 132 , pp. 333–341, 2020. doi: 10.1016/j.neunet.2020.09.010 32977278
12 Gao J. , and Xu C. S . CI-GNN: Building a category-instance graph for zero-shot video classification, IEEE Transactions on Multimedia, vol. 22 , no. 12 , pp. 3088–3100, 2020. doi: 10.1109/TMM.2020.2969787
13 Liu, S. C., Long, M. S., Wang, J. M., and Jordan, M. I. Generalized zero-shot learning with deep calibration network, Advances in Neural Information Processing Systems, vol. 31, 2018.
14 Sankaranarayanan, S., and Balaji, Y. Meta learning for domain generalization, In Meta Learning With Medical Imaging and Health Informatics Applications, pp. 75–86, 2023.
15 Xian, Y. Q., Lorenz, T., Schiele, B., and Akata, Z. Feature generating networks for zero-shot learning, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5542–5551, 2018.
16 Ren, J. W., Yu, C. J., Ma, X., Zhao, H. Y., Yi, S., et al. Balanced meta-softmax for long-tailed visual recognition, Advances in Neural Information Processing Systems, vol. 33, pp. 4175–4186, 2020.
17 Liu, M., Li, F., Zhang, C., Wei, Y., Bai, H., and Zhao, Y. Progressive semantic-visual mutual adaption for generalized zero-shot learning, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15337–15346, 2023.
18 Gao Y. , Feng W. , Xiao R. , He L. , He Z. , Lv J. , et al . Improving generalized zero-shot learning via cluster-based semantic disentangling representation, Pattern Recognition, vol. 150 , pp. 110320, 2024, Elsevier. doi: 10.1016/j.patcog.2024.110320
19 Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., et al. Learning transferable visual models from natural language supervision, In International Conference on Machine Learning, pp. 8748–8763, 2021.
20 Shipard, J., Wiliem, A., Thanh, K. N., Xiang, W., and Fookes, C. Diversity is definitely needed: Improving model-agnostic zero-shot classification via stable diffusion, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 769–778, 2023.
21 Christensen, A., Mancini, M., Koepke, A., Winther, O., and Akata, Z. Image-free classifier injection for zero-shot classification, Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 19072–19081, 2023.
22 Novack, Z., McAuley, J., Lipton, Z. C., and Garg, S. CHILS: Zero-shot image classification with hierarchical label sets, In International Conference on Machine Learning, pp. 26342–26362, 2023.
23 Zhou, Z., Lei, Y., Zhang, B., Liu, L., and Liu, Y. Zegclip: Towards adapting clip for zero-shot semantic segmentation, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11175–11185, 2023.
24 Guo, Z., Zhang, R., Qiu, L., Ma, X., Miao, X., He, X., et al. Calip: Zero-shot enhancement of clip with parameter-free attention, Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 1, pp. 746–754, 2023.
25 Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., et al. Emerging properties in self-supervised vision transformers, Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9650–9660, 2021.
26 Shakouri, M., Iranmanesh, F., and Eftekhari, M. DINO-CXR: A self-supervised method based on vision transformer for chest X-ray classification, In International Symposium on Visual Computing, pp. 320–331, 2023.
27 Li, F., Zhang, H., Xu, H., Liu, S., Zhang, L., Ni, L. M., et al. Mask dino: Towards a unified transformer-based framework for object detection and segmentation, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3041–3050, 2023.
28 Zhang, R., Hu, X., Li, B., Huang, S., Deng, H., Qiao, Y., et al. Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15211–15222, 2023.
29 Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. Improving language understanding by generative pre-training, San Francisco, CA, USA, 2018.
30 Baktash, J. A., and Dawodi, M. GPT-4: A review on advancements and opportunities in natural language processing, arXiv preprint arXiv:2305.03195, 2023
31 Reddy M. D. M. , Basha M. S. M. , Hari M. C. , and Penchalaiah N . Dall-e: Creating images from text, UGC Care Group I Journal, vol. 8 , no. 14 , pp. 71–75, 2021.
32 Bendale, A. and Boult, T. E., Towards open set deep networks, Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1563–1572, 2016.
33 Liang, S. and Li, Y. and Srikant, R., Enhancing the reliability of out-of-distribution image detection in neural networks, arXiv preprint arXiv:1706.02690, 2017.
34 Zhang, H. and Li, A. and Guo, J. and Guo, Y., Hybrid Models for Open Set Recognition, Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part III, pp. 102–117, Springer, 2020.
35 Chen, G. and Qiao, L. and Shi, Y. and Peng, P. and Li, J. and Huang, T. et al. Learning Open Set Network with Discriminative Reciprocal Points, Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part III, pp. 507–522, Springer, 2020.
36 Lu, J. and Xu, Y. and Li, H. and Cheng, Z. and Niu, Y., PMAL: Open Set Recognition via Robust Prototype Mining, Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 2, pp. 1872–1880, 2022.
37 Esmaeilpour, S. and Liu, B. and Robertson, E. and Shu, L., Zero-shot Out-of-Distribution Detection Based on the Pre-trained Model CLIP, Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 6, pp. 6568–6576, 2022.
38 Vaze S . and Han K. and Vedaldi A. and Zisserman A. , Open-Set Recognition: A Good Closed-Set Classifier Is All You Need?, Journal of Interesting Stuff, OpenReview, 2021.
39 Moon, W. and Park, J. and Seong, H. S. and Cho, C.-H. and Heo, J.-P., Difficulty-Aware Simulator for Open Set Recognition, European Conference on Computer Vision, pp. 365–381, Springer, 2022.
40 Cho, W. and Choo, J., Towards Accurate Open-Set Recognition via Background-Class Regularization, European Conference on Computer Vision, pp. 658–674, Springer, 2022.
41 Liu Z.-g . and Fu Y.-m . and Pan Q . and Zhang Z.-w. , Orientational Distribution Learning with Hierarchical Spatial Attention for Open Set Recognition, IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022. doi: 10.1109/TPAMI.2022.3227913
42 Huang H . and Wang Y . and Hu Q . and Cheng M.-M. , Class-Specific Semantic Reconstruction for Open Set Recognition, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45 , no. 4 , pp. 4214–4228, 2022.
