WUDI-FeatCal CLIP ViT-B/32 TA8

This repository contains a merged CLIP vision encoder for the TA8 image-classification benchmark. The model is based on openai/clip-vit-base-patch32 and was produced with WUDI-Merging followed by FeatCal calibration. This model is released as part of our work:

This is a transformers.CLIPVisionModel, not a full CLIPModel. Use the processor from openai/clip-vit-base-patch32.

FusionBench:

Load

from transformers import CLIPProcessor, CLIPVisionModel

model = CLIPVisionModel.from_pretrained("yanggangu/wudi-featcal-b32-ta8")
processor = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32")

Evaluation

Evaluation was run with FusionBench using:

fusion_bench \
  method=dummy \
  modelpool=CLIPVisionModelPool/clip-vit-base-patch32_TA8 \
  modelpool.models._pretrained_=yanggangu/wudi-featcal-b32-ta8 \
  taskpool=CLIPVisionModelTaskPool/clip-vit-classification_TA8 \
  seed=42
SUN397 Cars RESISC45 EuroSAT SVHN GTSRB MNIST DTD Average
72.2 76.0 93.0 98.1 96.7 97.9 99.4 77.1 88.8

Raw accuracies:

sun397:        0.7221158742904663
stanford-cars: 0.7598557472229004
resisc45:      0.9301587343215942
eurosat:       0.9807407259941101
svhn:          0.9672326445579529
gtsrb:         0.9792557358741760
mnist:         0.9944999814033508
dtd:           0.7707446813583374
average:       0.8880755156278610

Comparison With FusionBench Reference Results

The table below compares this model with selected CLIP ViT-B/32 TA8 reference results reported in the FusionBench documentation.

This model is not the overall best number in the public FusionBench table because Weight-Ensembling MoE reports a higher average score, but it is competitive with the strongest entries while remaining a single dense CLIPVisionModel with the same parameter count as the base vision encoder.

Method SUN397 Cars RESISC45 EuroSAT SVHN GTSRB MNIST DTD Average
Fine-tuned (STL) 75.0 78.3 95.2 99.0 97.3 98.9 99.6 79.7 90.3
Traditional MTL 72.3 76.6 92.2 97.9 95.5 97.7 99.3 77.7 88.6
RegMean++ 69.3 70.5 86.7 96.1 94.1 90.4 99.0 68.7 84.4
WUDI-Merging 68.0 72.5 85.0 94.6 94.8 94.9 99.3 66.6 84.5
Weight-Ensembling MoE 73.7 76.8 93.4 98.2 96.8 98.2 99.6 76.6 89.2
WUDI-Merging + FeatCal (this model) 72.2 76.0 93.0 98.1 96.7 97.9 99.4 77.1 88.8

Reference: https://tanganke.github.io/fusion_bench/modelpool/clip_vit/

Citation

@article{gu2026featcal,
  title={FeatCal: Feature Calibration for Post-Merging Models},
  author={Gu, Yanggan and Cai, Shuo and Wang, Zihao and Wang, Wenjun and Wang, Yuanyi and Wang, Pengkai and Huang, Sirui and Lu, Su and Wu, Jianmin and Yang, Hongxia},
  journal={arXiv preprint arXiv:2605.13030},
  year={2026}
}
Downloads last month
7
Safetensors
Model size
87.5M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yanggangu/wudi-featcal-b32-ta8

Finetuned
(155)
this model

Paper for yanggangu/wudi-featcal-b32-ta8