Instructions to use Etherll/Qwen3.6-27B-Layerdose with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Etherll/Qwen3.6-27B-Layerdose with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Etherll/Qwen3.6-27B-Layerdose") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Etherll/Qwen3.6-27B-Layerdose") model = AutoModelForMultimodalLM.from_pretrained("Etherll/Qwen3.6-27B-Layerdose", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Etherll/Qwen3.6-27B-Layerdose with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Etherll/Qwen3.6-27B-Layerdose" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Etherll/Qwen3.6-27B-Layerdose", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Etherll/Qwen3.6-27B-Layerdose
- SGLang
How to use Etherll/Qwen3.6-27B-Layerdose with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Etherll/Qwen3.6-27B-Layerdose" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Etherll/Qwen3.6-27B-Layerdose", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Etherll/Qwen3.6-27B-Layerdose" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Etherll/Qwen3.6-27B-Layerdose", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Etherll/Qwen3.6-27B-Layerdose with Docker Model Runner:
docker model run hf.co/Etherll/Qwen3.6-27B-Layerdose
Whats layerdose?
Hi,
Any more info about this?
Hi, this is just a personal LayerDose experiment, not a model Iโd recommend running yet.
The idea was to test whether some low-risk linear_attn operators in Qwen3.6-27B could be replaced with much smaller rank-64 patch modules. In plain terms: instead of keeping every original operator tensor, I replace selected low-impact operators with compact learned patches that try to approximate their effect. That cuts checkpoint storage and should reduce some compute, while hopefully keeping behavior close to the original model.
For this test build, the replaced tensors save about 3.28 GB before adding the small patch file, and the early speed result is around 10% faster. The small smoke checks looked okay, but this is not a full benchmark or a finished release. It also needs the LayerDose patch loader to activate the replacements, so a normal from_pretrained() load is not enough.
So short version: it is an experimental compression/replacement candidate. Interesting for research, but nothing worth depending on yet.