QUEEN (HCE-4)

Chess analysis and move selection with a fine-tuned SmolLM3 decoder and LC0 BT5 board encoder. All weights and inference code are included; no separate LC0 or Stockfish installation is needed. Use the runner below, not AutoModelForCausalLM alone (which omits board conditioning).

Run

Requires Linux, an NVIDIA GPU with BF16 support, a CUDA-13-compatible driver, and uv. uv downloads Python automatically. Model downloads total about 8.4 GiB; the CUDA environment needs additional disk space. Tested hardware is an H100 80GB; smaller GPUs are not yet validated.

uv venv --python 3.12.12 queen_hce-4/.venv
uv pip install --python queen_hce-4/.venv/bin/python huggingface-hub==1.27.0
queen_hce-4/.venv/bin/hf download princeton-nlp/queen_hce-4 --local-dir queen_hce-4
uv pip sync --python queen_hce-4/.venv/bin/python queen_hce-4/requirements-inference.txt
queen_hce-4/.venv/bin/python queen_hce-4/infer.py --fen 'rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1'

CUDA is detected automatically. To select a GPU, prefix the last command with CUDA_VISIBLE_DEVICES=0. CPU inference is not supported.

Defaults match this model's Elo evaluation: temperature 0.6, top-k 20, top-p 0.95, 2048 output tokens, repetition penalty 1.0, BF16, V1 eager execution, prefix caching off, context length 2304, and batch capacity 32. The model-specific evaluation prompt and move-selection policy are in inference_config.json. The base seed is 20260823; request seeds use the evaluation's hash of base seed, --game-id (default 0), and ply from the FEN. --history history.json supplies prior FENs in chronological order, excluding the current position. Supply the actual game history and game index to reproduce an evaluation request; a FEN alone cannot recover history. Exact sampled text can still vary with hardware, library versions, and batching.

Output is JSON after library logs: decoded text, raw_text, best_move_uci, move_source, a legality-checked critical line, and the actual inference settings. As in Elo evaluation, unparseable responses use a seeded random legal move; move_source explicitly labels this fallback. --gpu-memory-utilization defaults to 0.72; reduce it when sharing a GPU, leaving room for the separately loaded encoder. Changing generation settings overrides evaluation defaults.

Notes

Explanations and variations may be wrong; this is not a UCI executable. VALIDATION.json records tests; release.json contains file checksums. The pinned environment uses PyTorch 2.13.0, Transformers 5.15.0 and vLLM 0.27.1.

Base model: SmolLM3-3B (Apache-2.0). Encoder: Leela Chess Zero, BT5-1024x15x32h-rpe-swa-3700000. Upstream component terms apply; no blanket license for the combined fine-tuned release is asserted here.

Downloads last month
44
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for princeton-nlp/queen_hce-4

Finetuned
(163)
this model

Collection including princeton-nlp/queen_hce-4