Instructions to use yanggangu/wudi-featcal-b32-ta8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use yanggangu/wudi-featcal-b32-ta8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-classification", model="yanggangu/wudi-featcal-b32-ta8") pipe("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/hub/parrots.png")# Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("yanggangu/wudi-featcal-b32-ta8") model = AutoModel.from_pretrained("yanggangu/wudi-featcal-b32-ta8", device_map="auto") - Notebooks
- Google Colab
- Kaggle
WUDI-FeatCal CLIP ViT-B/32 TA8
This repository contains a merged CLIP vision encoder for the TA8 image-classification benchmark.
The model is based on openai/clip-vit-base-patch32 and was produced with WUDI-Merging followed by FeatCal calibration.
This model is released as part of our work:
- FeatCal: Feature Calibration for Post-Merging Models
- Paper: https://arxiv.org/abs/2605.13030
This is a transformers.CLIPVisionModel, not a full CLIPModel. Use the processor from openai/clip-vit-base-patch32.
FusionBench:
- GitHub: https://github.com/tanganke/fusion_bench
- Documentation: https://tanganke.github.io/fusion_bench/
Load
from transformers import CLIPProcessor, CLIPVisionModel
model = CLIPVisionModel.from_pretrained("yanggangu/wudi-featcal-b32-ta8")
processor = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32")
Evaluation
Evaluation was run with FusionBench using:
fusion_bench \
method=dummy \
modelpool=CLIPVisionModelPool/clip-vit-base-patch32_TA8 \
modelpool.models._pretrained_=yanggangu/wudi-featcal-b32-ta8 \
taskpool=CLIPVisionModelTaskPool/clip-vit-classification_TA8 \
seed=42
| SUN397 | Cars | RESISC45 | EuroSAT | SVHN | GTSRB | MNIST | DTD | Average |
|---|---|---|---|---|---|---|---|---|
| 72.2 | 76.0 | 93.0 | 98.1 | 96.7 | 97.9 | 99.4 | 77.1 | 88.8 |
Raw accuracies:
sun397: 0.7221158742904663
stanford-cars: 0.7598557472229004
resisc45: 0.9301587343215942
eurosat: 0.9807407259941101
svhn: 0.9672326445579529
gtsrb: 0.9792557358741760
mnist: 0.9944999814033508
dtd: 0.7707446813583374
average: 0.8880755156278610
Comparison With FusionBench Reference Results
The table below compares this model with selected CLIP ViT-B/32 TA8 reference results reported in the FusionBench documentation.
This model is not the overall best number in the public FusionBench table because Weight-Ensembling MoE reports a higher average score, but it is competitive with the strongest entries while remaining a single dense CLIPVisionModel with the same parameter count as the base vision encoder.
| Method | SUN397 | Cars | RESISC45 | EuroSAT | SVHN | GTSRB | MNIST | DTD | Average |
|---|---|---|---|---|---|---|---|---|---|
| Fine-tuned (STL) | 75.0 | 78.3 | 95.2 | 99.0 | 97.3 | 98.9 | 99.6 | 79.7 | 90.3 |
| Traditional MTL | 72.3 | 76.6 | 92.2 | 97.9 | 95.5 | 97.7 | 99.3 | 77.7 | 88.6 |
| RegMean++ | 69.3 | 70.5 | 86.7 | 96.1 | 94.1 | 90.4 | 99.0 | 68.7 | 84.4 |
| WUDI-Merging | 68.0 | 72.5 | 85.0 | 94.6 | 94.8 | 94.9 | 99.3 | 66.6 | 84.5 |
| Weight-Ensembling MoE | 73.7 | 76.8 | 93.4 | 98.2 | 96.8 | 98.2 | 99.6 | 76.6 | 89.2 |
| WUDI-Merging + FeatCal (this model) | 72.2 | 76.0 | 93.0 | 98.1 | 96.7 | 97.9 | 99.4 | 77.1 | 88.8 |
Reference: https://tanganke.github.io/fusion_bench/modelpool/clip_vit/
Citation
@article{gu2026featcal,
title={FeatCal: Feature Calibration for Post-Merging Models},
author={Gu, Yanggan and Cai, Shuo and Wang, Zihao and Wang, Wenjun and Wang, Yuanyi and Wang, Pengkai and Huang, Sirui and Lu, Su and Wu, Jianmin and Yang, Hongxia},
journal={arXiv preprint arXiv:2605.13030},
year={2026}
}
- Downloads last month
- 7
Model tree for yanggangu/wudi-featcal-b32-ta8
Base model
openai/clip-vit-base-patch32