ZAYA1-base โ€” GGUF

Our own GGUF conversion of Zyphra's ZAYA1-base. It keeps Zyphra's original Megatron-style checkpoint layout, which transformers cannot load; our converter maps it to ZAYA's standard layout (llama.cpp #14). Converting Zyphra/ZAYA1-8B-legacy this way gives tensors byte-identical to the ones from Zyphra/ZAYA1-8B.

Contents

  • ZAYA1-base-Q4_K_M.gguf (8B shape: 40 layers, 16 experts; rope theta 1e6; 32,768-token context).

Validation

  • Wikitext-2 test perplexity, 60 chunks of 512 tokens, Q4_K_M on Vulkan (Radeon 8060S): 8.53 ยฑ 0.17.
  • Raw-text perplexity favours base models; the post-trained ZAYA1-8B scores 32.13 on the same test.

Running it

With the 1bit engine:

1bit serve -m ZAYA1-base-Q4_K_M.gguf --device vulkan

ZAYA runs from our llama.cpp fork (branch 1bit/hrx-vulkan-patched); upstream llama.cpp has no ZAYA model. The GGUFs must come from our converter: it writes the grouped convolution's weights tap-major, which the graph expects.

Attribution

  • Base model: Zyphra/ZAYA1-base, Apache 2.0.
  • GGUF conversion and validation: the 1bit engine project, with the converter and model code in our llama.cpp fork.
  • License: Apache 2.0, inherited from the base model.
Downloads last month
144
GGUF
Model size
9B params
Architecture
zaya
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for 1bit-MONSTER/ZAYA1-base-GGUF

Quantized
(3)
this model