QUEEN (HCE-4)
Chess analysis and move selection with a fine-tuned SmolLM3 decoder and LC0 BT5
board encoder. All weights and inference code are included; no separate LC0 or
Stockfish installation is needed. Use the runner below, not AutoModelForCausalLM
alone (which omits board conditioning).
Run
Requires Linux, an NVIDIA GPU with BF16 support, a CUDA-13-compatible driver, and uv. uv downloads Python automatically. Model downloads total about 8.4 GiB; the CUDA environment needs additional disk space. Tested hardware is an H100 80GB; smaller GPUs are not yet validated.
uv venv --python 3.12.12 queen_hce-4/.venv
uv pip install --python queen_hce-4/.venv/bin/python huggingface-hub==1.27.0
queen_hce-4/.venv/bin/hf download princeton-nlp/queen_hce-4 --local-dir queen_hce-4
uv pip sync --python queen_hce-4/.venv/bin/python queen_hce-4/requirements-inference.txt
queen_hce-4/.venv/bin/python queen_hce-4/infer.py --fen 'rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1'
CUDA is detected automatically. To select a GPU, prefix the last command with
CUDA_VISIBLE_DEVICES=0. CPU inference is not supported.
Defaults match this model's Elo evaluation: temperature 0.6, top-k 20,
top-p 0.95, 2048 output tokens, repetition penalty 1.0, BF16, V1 eager execution,
prefix caching off, context length 2304, and batch capacity 32.
The model-specific evaluation prompt and move-selection policy are in
inference_config.json. The base seed is 20260823; request seeds use the evaluation's
hash of base seed, --game-id (default 0), and ply from the FEN.
--history history.json supplies prior FENs in chronological order, excluding
the current position. Supply the actual game history and game index to reproduce
an evaluation request; a FEN alone cannot recover history. Exact sampled text
can still vary with hardware, library versions, and batching.
Output is JSON after library logs: decoded text, raw_text, best_move_uci,
move_source, a legality-checked critical line, and the actual inference settings.
As in Elo evaluation, unparseable responses use a seeded random legal move;
move_source explicitly labels this fallback. --gpu-memory-utilization defaults
to 0.72; reduce it when sharing a GPU, leaving room for the separately loaded encoder.
Changing generation settings overrides evaluation defaults.
Notes
Explanations and variations may be wrong; this is not a UCI executable.
VALIDATION.json records tests; release.json contains file checksums.
The pinned environment uses PyTorch 2.13.0, Transformers 5.15.0 and vLLM 0.27.1.
Base model: SmolLM3-3B (Apache-2.0). Encoder: Leela Chess Zero, BT5-1024x15x32h-rpe-swa-3700000. Upstream component terms apply; no blanket license for the combined fine-tuned release is asserted here.
- Downloads last month
- 44
Model tree for princeton-nlp/queen_hce-4
Base model
HuggingFaceTB/SmolLM3-3B-Base