YOLOS-Tiny for Retail Shelf Object Detection

This repository contains a fine-tuned version of the hustvl/yolos-tiny architecture, optimized specifically for dense object detection on retail store shelves using the SKU-110k dataset layout.

Model Description

The core model is based on the You Only Look at One Sequence object detection transformer. The architecture has been adapted to handle high-density patterns of closely packed objects, typical in retail environments (shelf auditing and planogram compliance analysis).

  • Developer: Bartek
  • Model Type: Object Detection Transformer (ViT-based)
  • Base Model: hustvl/yolos-tiny
  • Language: Python / PyTorch

Training Hyperparameters

The model was optimized using the full training split of the target dataset with the following pipeline parameters:

  • Optimizer: AdamW
  • Learning Rate: 5e-5 (initial)
  • Learning Rate Scheduler: CosineAnnealingLR
  • Batch Size: 4
  • Epochs: 5
  • Data Pipeline: Custom sparse collate function handling mixed dimension object boundaries.

Intended Uses & Limitations

This model is intended for custom shelf analytics deployment, automation of out-of-stock monitoring, and object extraction workflows.

How to use:

from transformers import YolosForObjectDetection, YolosImageProcessor
import torch

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

model = YolosForObjectDetection.from_pretrained("Bartek7630/yolos-tiny-sku110k-refined").to(device)
image_processor = YolosImageProcessor.from_pretrained("Bartek7630/yolos-tiny-sku110k-refined")

### Full Inference & Visualization Script

You can copy and run the complete pipeline below to test the model on any retail shelf image.

```python
import torch
import requests
from PIL import Image, ImageDraw
from transformers import YolosForObjectDetection, YolosImageProcessor

# 1. Configuration & Environment Setup
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model_id = "Bartek7630/yolos-tiny-sku110k-refined"

# 2. Load Model and Processor directly from Hugging Face Hub
model = YolosForObjectDetection.from_pretrained(model_id).to(device)
image_processor = YolosImageProcessor.from_pretrained(model_id)
model.eval()

# 3. Load Input Image (Replace URL with local path if necessary)
url = "[https://raw.githubusercontent.com/huggingface/transformers/main/tests/fixtures/tests_samples/COCO/000000039769.png](https://raw.githubusercontent.com/huggingface/transformers/main/tests/fixtures/tests_samples/COCO/000000039769.png)"
image = Image.open(requests.get(url, stream=True).raw).convert("RGB")

# 4. Preprocessing
inputs = image_processor(images=image, return_tensors="pt").to(device)

# 5. Model Inference
with torch.no_grad():
    outputs = model(**inputs)

# 6. Post-Processing (Bounding Box Denormalization)
target_sizes = torch.tensor([image.size[::-1]]).to(device)
results = image_processor.post_process_object_detection(outputs, threshold=0.3, target_sizes=target_sizes)[0]

# 7. Visualization
draw = ImageDraw.Draw(image)
for score, label, box in zip(results["scores"], results["labels"], results["boxes"]):
    box = [int(i) for i in box.tolist()]
    
    # Draw red rectangles over detected SKU items
    draw.rectangle(box, outline="red", width=3)

# Save or display the finalized audit result
image.save("shelf_audit_result.png")
image.show()
Downloads last month
15
Safetensors
Model size
6.47M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Bartek7630/yolos-tiny-sku110k-refined

Space using Bartek7630/yolos-tiny-sku110k-refined 1