Instructions to use deepseek-ai/DeepSeek-V2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use deepseek-ai/DeepSeek-V2 with Transformers:

# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="deepseek-ai/DeepSeek-V2", trust_remote_code=True)
messages = [
    {"role": "user", "content": "Who are you?"},
]
pipe(messages)

# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V2", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("deepseek-ai/DeepSeek-V2", trust_remote_code=True)
messages = [
    {"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=40)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))

Notebooks
Google Colab
Kaggle
Local Apps Settings

vLLM

How to use deepseek-ai/DeepSeek-V2 with vLLM:

Install from pip and serve model

# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "deepseek-ai/DeepSeek-V2"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "deepseek-ai/DeepSeek-V2",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Use Docker

docker model run hf.co/deepseek-ai/DeepSeek-V2

SGLang

How to use deepseek-ai/DeepSeek-V2 with SGLang:

Install from pip and serve model

# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "deepseek-ai/DeepSeek-V2" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "deepseek-ai/DeepSeek-V2",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Use Docker images

docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "deepseek-ai/DeepSeek-V2" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "deepseek-ai/DeepSeek-V2",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Docker Model Runner
How to use deepseek-ai/DeepSeek-V2 with Docker Model Runner:
```
docker model run hf.co/deepseek-ai/DeepSeek-V2
```

KV Cache for compress_kv or key-value states

by House-99 - opened May 7, 2024

Discussion

House-99

May 7, 2024

From tech report, when it comes to inference, the kv cache should be compress_kv. But from the modeling_deepseek.py, I notice that the key state and value state still like llama2. Something wrongs right?

Xidong

May 7, 2024

This comment has been hidden

ZhongYingMatrix

May 7, 2024

Does the file modeling_deepseek.py merely contain executable code accompanying open-source weights, without the actual implementation of compressed kv? I've also noticed that the training implementation for DeepseekV2MoE lacks supporting of ep.

msr2000

DeepSeek org May 7, 2024

Thank you for your interest in our work. We are aware of the challenges in implementing KV compression on current open-source code and are actively working on it. The HuggingFace's code is not as efficient as we would like, so we're developing a new open-source code using vLLM for better performance. The open-source vLLM code including KV compression will be released once it is ready.

House-99

May 7, 2024

Does the file modeling_deepseek.py merely contain executable code accompanying open-source weights, without the actual implementation of compressed kv? I've also noticed that the training implementation for DeepseekV2MoE lacks supporting of ep.

DeepseekV2 opensource their training implementation? I do not find the link yet.

ZhongYingMatrix

May 7, 2024

DeepseekV2 opensource their training implementation? I do not find the link yet.

Apologies for the confusion, I mean the implementation in modeling_deepseek.py

Jiayi-Pan

Jun 25, 2024

•

edited Jun 25, 2024

@msr2000 Thank you for your contribution to the OS community.
I am working on a project where we'd like to use DS-2 as the base model. Inference speed has become a bottleneck for our project
We wonder if you have any etas on the timeline for open sourcing efficient vllm inference code?

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment