Instructions to use ornith-ai/Ornith-1.5-35B-A3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ornith-ai/Ornith-1.5-35B-A3B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ornith-ai/Ornith-1.5-35B-A3B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ornith-ai/Ornith-1.5-35B-A3B") model = AutoModelForMultimodalLM.from_pretrained("ornith-ai/Ornith-1.5-35B-A3B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ornith-ai/Ornith-1.5-35B-A3B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ornith-ai/Ornith-1.5-35B-A3B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ornith-ai/Ornith-1.5-35B-A3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ornith-ai/Ornith-1.5-35B-A3B
- SGLang
How to use ornith-ai/Ornith-1.5-35B-A3B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ornith-ai/Ornith-1.5-35B-A3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ornith-ai/Ornith-1.5-35B-A3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ornith-ai/Ornith-1.5-35B-A3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ornith-ai/Ornith-1.5-35B-A3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ornith-ai/Ornith-1.5-35B-A3B with Docker Model Runner:
docker model run hf.co/ornith-ai/Ornith-1.5-35B-A3B
No speculative decoding, No QAT
Hi team
Please look at other teams like LFM, Google gemma. Everyone shipping MTP, or Dflash or Dspark for speculative decoding
Also LFM is releasing QAD model that is 97% accurate to BF16 with dspark speculative decoding for 3B model. Similarly gemma has MTP with QAT releases.
I hope you guys take inspiration from that.
The companies you mention actually make LLMs, ornith-ai so far only make finetunes of existing LLMs, they re-trained existing LLMs with their own datasets or whatever, however the MTPs do not work correctly because they were made by Qwen for the oiriginal/official model, not for the onith-ai finetune(s).
MTP heads are not that hard to train (and can and probably should be done during the tuning stage) - luckily it's also easy to do afterwards (mine was done in less than a day: https://huggingface.co/shisa-ai/Ornith-1.5-35B-A3B-MTP-ONLY ).
DFlash does takes more effort (days to get a decent version). QAT is also practically free while training, but the QAT is targeted for a single specific quant, so if home users are mostly going to use various GGUF quants, it doesn't much matter.
MTP heads are not that hard to train (and can and probably should be done during the tuning stage) - luckily it's also easy to do afterwards (mine was done in less than a day: https://huggingface.co/shisa-ai/Ornith-1.5-35B-A3B-MTP-ONLY ).
DFlash does takes more effort (days to get a decent version).
True that is the point. They can if they are willing too.
QAT is also practically free while training, but the QAT is targeted for a single specific quant, so if home users are mostly going to use various GGUF quants, it doesn't much matter.
QAT or QAD reduces error when it is quantized to Q4. Which is not same as quants when QAT or QAD is not done. Take LFM example
I think the more important point is QAT has to be targeted for a specific target, eg NVFP4, MXFP4, GPTQ W4A16, Q4_0, etc. This is only useful for someone running an "official" quant for that specific target/layout. In practice, almost all "local" users are running GGUF quants. A Q4_K is actually a mix of Q4_K, Q5_K, Q6_K and Q8_0 (and some F32) tensors.
Note: as Unsloth showed w/ the Gemma QAT models even a Q4_0 safetensor -> Q4_0 GGUF does not protect accuracy/causes severe degradation unless you do additional work.
Gemma team has provided Q4_0 and same for LFM Q4_0. Q4_0 is easy inference for all backends. Only gpt oss was MXFP4. That is also properly supported on all backends in llama.cpp
ornith-ai fixed the issue a day ago:
https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B/blob/main/model-00016-of-00016.safetensors
The companies you mention actually make LLMs, ornith-ai so far only make finetunes of existing LLMs, they re-trained existing LLMs with their own datasets or whatever, however the MTPs do not work correctly because they were made by Qwen for the oiriginal/official model, not for the onith-ai finetune(s).
Ironically the original Qwen3.6-35B MTP works because it's trained, while Ornith's is not. So your statement is wrong on both fact and reason.