Instructions to use yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5 with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5 with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Ornith-1.0-9B-oQe3.5
Apple Silicon Optimized oQe3.5 MLX Quantized Release
This repository contains an oQe3.5 mixed-precision MLX quantized version of Ornith-1.0-9B, optimized for fast and memory-efficient inference on Apple Silicon.
The original Ornith-1.0-9B model was developed by DeepReinforce. This repository provides an optimized MLX conversion only and does not contain any additional fine-tuning or retraining. The goal is to preserve the original model's exceptional coding and reasoning performance while significantly reducing memory requirements through oQe3.5 sensitivity-aware mixed-precision quantization. :contentReference[oaicite:0]{index=0}
About Ornith-1.0
Ornith-1.0 is a family of open-source reasoning models specialized for agentic software engineering and coding workflows.
The 9B model is the most compact member of the Ornith family, delivering strong performance while remaining practical for local deployment.
Key capabilities include:
- 🧠 Advanced reasoning
- 💻 Agentic coding
- 🛠 Native tool calling
- 🔧 Function calling
- 📚 Long-context reasoning
- 🤖 Multi-step planning
- 🔍 Codebase understanding
- ⚡ Software engineering automation
Ornith is trained using a reinforcement learning framework that jointly optimizes both solution generation and the reasoning scaffolds used to reach those solutions, enabling stronger search trajectories and higher-quality coding performance. :contentReference[oaicite:1]{index=1}
Quantization
This release uses oQe3.5 mixed-precision quantization.
Specifications
- Format: MLX
- Quantization: oQe3.5
- Method: Sensitivity-Aware Mixed Precision
- Platform: Apple Silicon
- Inference Engine: MLX / oMLX
Unlike traditional fixed-bit quantization, oQe3.5 dynamically assigns precision based on layer sensitivity, preserving higher precision where it matters most while aggressively compressing less sensitive layers.
Benefits include:
- Higher reasoning quality
- Better coding accuracy
- Lower memory usage
- Faster inference
- Excellent Apple Silicon efficiency
Recommended Settings
For the best overall performance:
temp: 0.6
top_p: 0.95
top_k: 20
min_p: 0
rep_penalty: 1.05
presence_penalty: 1.2
enable_thinking: true
For benchmark reproduction:
temp: 1.0
top_p: 0.95
top_k: 20
These settings align with the recommendations from the original Ornith model card. :contentReference[oaicite:2]{index=2}
Example Usage
from mlx_lm import load, generate
model, tokenizer = load("yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5")
messages = [
{
"role": "user",
"content": "Write a Python implementation of an LRU cache."
}
]
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=True,
)
response = generate(
model,
tokenizer,
prompt=prompt,
temp=0.6,
top_p=0.95,
top_k=20,
max_tokens=16384,
)
print(response)
Optimized For
This release is optimized for:
- Apple M1
- Apple M2
- Apple M3
- Apple M4
Compatible with:
- MLX
- MLX-LM
- oMLX
- Open WebUI
- LM Studio (MLX)
- Local AI Agent Frameworks
Model Highlights
- State-of-the-art open coding model in the 9B class
- Advanced reasoning
- Agentic software engineering
- Native function calling
- Native tool use
- Long-context reasoning
- Multi-step planning
- Repository understanding
- Research assistance
Intended Use
This model is particularly well suited for:
- Software Engineering
- AI Coding Assistants
- Autonomous Coding Agents
- Repository Analysis
- Bug Fixing
- Code Generation
- Research Automation
- Multi-step Planning
- Local AI Development
Hardware Recommendations
Recommended systems:
- Apple M1 Pro / Max / Ultra
- Apple M2 Pro / Max / Ultra
- Apple M3 Series
- Apple M4 Series
Higher-memory Apple Silicon systems provide the best experience for long-context coding sessions.
About oQe3.5 Quantization
oQe3.5 is a sensitivity-aware mixed-precision quantization method developed to maximize model quality while substantially reducing memory requirements.
Instead of assigning the same precision to every weight matrix, oQe3.5 automatically allocates precision according to each layer's importance.
This approach enables:
- Better reasoning preservation
- Higher coding quality
- Faster inference
- Lower RAM usage
- Excellent Apple Silicon performance
Original Model
The original Ornith-1.0-9B introduces a self-improving reinforcement learning framework for agentic coding and is released under the MIT License.
Notable features include:
- Reinforcement Learning for agentic reasoning
- Tool-aware reasoning
- Native XML tool calling
- Coding-first optimization
- 262K context window
- Open-source MIT license
For detailed benchmarks, evaluation methodology, and technical documentation, please visit the original model page. :contentReference[oaicite:3]{index=3}
Credits
Original Model
All credit for the original model, datasets, training methodology, evaluation, benchmarks, and research belongs entirely to:
DeepReinforce Team
Original Repository:
https://huggingface.co/deepreinforce-ai/Ornith-1.0-9B
Project Page:
https://deep-reinforce.com/ornith_1_0.html
oQe3.5 MLX Quantized Release
This repository provides an Apple Silicon optimized oQe3.5 MLX quantized version of the original model.
No additional fine-tuning has been performed.
Acknowledgements
- DeepReinforce
- Qwen Team
- Apple MLX
- Hugging Face
- Transformers
- vLLM
- SGLang
- oMLX
- OptiQ Quantization
Citation
If you use this model in research, please cite the original Ornith paper:
@misc{ornith_9b,
title={{Ornith-1.0-9B}: Agentic Coding, Open to All},
author={DeepReinforce Team},
year={2026},
url={https://deep-reinforce.com/ornith_1_0.html}
}
License
This release inherits the MIT License from the original model.
Please refer to the original repository for complete licensing information.
Disclaimer
This repository contains an optimized oQe3.5 MLX quantized conversion intended for efficient local inference on Apple Silicon.
All original model architecture, datasets, training methodology, benchmarks, evaluations, and research remain entirely the work of the original DeepReinforce team.
- Downloads last month
- 34
3-bit
Model tree for yugeshkarunamurthy/Ornith-1.0-9B-oQe3.5
Base model
ornith-ai/Ornith-1.0-9B