Text Generation
Transformers
Safetensors
Serbian
mistral
mergekit
Merge
text-generation-inference
conversational
Instructions to use datatab/Yugo55A-GPT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use datatab/Yugo55A-GPT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="datatab/Yugo55A-GPT") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("datatab/Yugo55A-GPT") model = AutoModelForCausalLM.from_pretrained("datatab/Yugo55A-GPT", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use datatab/Yugo55A-GPT with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "datatab/Yugo55A-GPT" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "datatab/Yugo55A-GPT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/datatab/Yugo55A-GPT
- SGLang
How to use datatab/Yugo55A-GPT with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "datatab/Yugo55A-GPT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "datatab/Yugo55A-GPT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "datatab/Yugo55A-GPT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "datatab/Yugo55A-GPT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use datatab/Yugo55A-GPT with Docker Model Runner:
docker model run hf.co/datatab/Yugo55A-GPT
| base_model: | |
| - mlabonne/AlphaMonarch-7B | |
| - datatab/Yugo55-GPT-v4 | |
| - datatab/Yugo55-GPT-DPO-v1-chkp-300 | |
| - NousResearch/Nous-Hermes-2-Mistral-7B-DPO | |
| library_name: transformers | |
| tags: | |
| - mergekit | |
| - merge | |
| - text-generation-inference | |
| - transformers | |
| - mistral | |
| license: mit | |
| language: | |
| - sr | |
| datasets: | |
| - datatab/alpaca-cleaned-serbian-full | |
| - datatab/ultrafeedback_binarized | |
| - datatab/open-orca-slim-serbian | |
| # Yugo55A-GPT | |
| - **Developed by:** datatab | |
| - **License:** mit | |
| ## 🏆 Results | |
| > Results obtained through the Serbian LLM evaluation, released by Aleksa Gordić: [serbian-llm-eval](https://github.com/gordicaleksa/serbian-llm-eval) | |
| > * Evaluation was conducted on a 4-bit version of the model due to hardware resource constraints. | |
| <table> | |
| <tr> | |
| <th>MODEL</th> | |
| <th>ARC-E</th> | |
| <th>ARC-C</th> | |
| <th>Hellaswag</th> | |
| <th>BoolQ</th> | |
| <th>Winogrande</th> | |
| <th>OpenbookQA</th> | |
| <th>PiQA</th> | |
| </tr> | |
| <tr> | |
| <td><a href="https://huggingface.co/datatab/Yugo55-GPT-v4-4bit/">*Yugo55-GPT-v4-4bit</a></td> | |
| <td>51.41</td> | |
| <td>36.00</td> | |
| <td>57.51</td> | |
| <td>80.92</td> | |
| <td><strong>65.75</strong></td> | |
| <td>34.70</td> | |
| <td><strong>70.54</strong></td> | |
| </tr> | |
| <tr> | |
| <td><a href="https://huggingface.co/datatab/Yugo55A-GPT/">Yugo55A-GPT</a></td> | |
| <td><strong>51.52</strong></td> | |
| <td><strong>37.78</strong></td> | |
| <td><strong>57.52</strong></td> | |
| <td><strong>84.40</strong></td> | |
| <td>65.43</td> | |
| <td><strong>35.60</strong></td> | |
| <td>69.43</td> | |
| </tr> | |
| </table> | |
| # 🔗 Merge Details | |
| ### Merge Method | |
| > This is a merge of pre-trained language models created using [mergekit](https://github.com/cg123/mergekit). | |
| > This model was merged using the [linear](https://arxiv.org/abs/2203.05482) merge method. | |
| ### Models Merged | |
| The following models were included in the merge: | |
| * [datatab/Yugo55-GPT-v4](https://huggingface.co/datatab/Yugo55-GPT-v4) | |
| * [datatab/Yugo55-GPT-DPO-v1-chkp-300](https://huggingface.co/datatab/Yugo55-GPT-DPO-v1-chkp-300) | |
| * [mlabonne/AlphaMonarch-7B](https://huggingface.co/mlabonne/AlphaMonarch-7B) | |
| * [NousResearch/Nous-Hermes-2-Mistral-7B-DPO](https://huggingface.co/NousResearch/Nous-Hermes-2-Mistral-7B-DPO) | |
| ## 🧩 Configuration | |
| The following YAML configuration was used to produce this model: | |
| ```yaml | |
| models: | |
| - model: datatab/Yugo55-GPT-v4 | |
| parameters: | |
| weight: 1.0 | |
| - model: datatab/Yugo55-GPT-DPO-v1-chkp-300 | |
| parameters: | |
| weight: 1.0 | |
| - model: mlabonne/AlphaMonarch-7B | |
| parameters: | |
| weight: 0.5 | |
| - model: NousResearch/Nous-Hermes-2-Mistral-7B-DPO | |
| parameters: | |
| weight: 0.5 | |
| merge_method: linear | |
| dtype: float16 | |
| ``` | |
| ## 💻 Usage | |
| ```terminal | |
| !pip -q install git+https://github.com/huggingface/transformers # need to install from github | |
| !pip install -q datasets loralib sentencepiece | |
| !pip -q install bitsandbytes accelerate | |
| ``` | |
| ```python | |
| from IPython.display import HTML, display | |
| def set_css(): | |
| display(HTML(''' | |
| <style> | |
| pre { | |
| white-space: pre-wrap; | |
| } | |
| </style> | |
| ''')) | |
| get_ipython().events.register('pre_run_cell', set_css) | |
| ``` | |
| ```python | |
| import torch | |
| import transformers | |
| from transformers import AutoTokenizer, AutoModelForCausalLM | |
| model = AutoModelForCausalLM.from_pretrained( | |
| "datatab/Yugo55A-GPT", torch_dtype="auto" | |
| ) | |
| tokenizer = AutoTokenizer.from_pretrained( | |
| "datatab/Yugo55A-GPT", torch_dtype="auto" | |
| ) | |
| ``` | |
| ```python | |
| from typing import Optional | |
| from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer | |
| def generate( | |
| user_content: str, system_content: Optional[str] = "" | |
| ) -> str: | |
| system_content = "Ispod je uputstvo koje opisuje zadatak, upareno sa unosom koji pruža dodatni kontekst. Napišite odgovor koji na odgovarajući način kompletira zahtev." | |
| messages = [ | |
| { | |
| "role": "system", | |
| "content": system_content, | |
| }, | |
| {"role": "user", "content": user_content}, | |
| ] | |
| tokenized_chat = tokenizer.apply_chat_template( | |
| messages, tokenize=True, add_generation_prompt=True, return_tensors="pt" | |
| ).to("cuda") | |
| text_streamer = TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True) | |
| output = model.generate( | |
| tokenized_chat, | |
| streamer=text_streamer, | |
| max_new_tokens=2048, | |
| temperature=0.1, | |
| repetition_penalty=1.11, | |
| top_p=0.92, | |
| top_k=1000, | |
| pad_token_id=tokenizer.pad_token_id, | |
| eos_token_id=tokenizer.eos_token_id, | |
| do_sample=True, | |
| ) | |
| generated_text = tokenizer.decode(output[0], skip_special_tokens=True) | |
| ``` | |
| ```python | |
| generate("Nabroj mi sve planete suncevog sistemai reci mi koja je najveca planeta") | |
| ``` | |
| ```python | |
| generate("Koja je razlika između lame, vikune i alpake?") | |
| ``` | |
| ```python | |
| generate("Napišite kratku e-poruku Semu Altmanu dajući razloge za GPT-4 otvorenog koda") | |
| ``` |