Anima CJK vocab pack (preview)
Settings for this image
anima-base-v1.0, anima_cjk_vocab_pack_preview3 (per-glyph routing on; ComfyUI needs ComfyUI-Anima_lora-Adapter ≥ 3.12.0), no LoRA, seed 72, 30 steps, CFG 4, 768×1344, then ResShift ×2 upscale to 1536×2688.
prompt: sensitive, 1girl, chitanda eru, hyouka, @channel \(caststation\), looking at viewer, solo, photo background, indoors, skirt, speech bubble, japanese text. She is at the bar. She is saying "へんたい".
negative: simple background, greyscale, sepia, blurry, glitch, jpeg artifacts, bad hands, bad anatomy, multiple views, thick border, looking from behind
ComfyUI: load the pack with the AnimaVocabPackLoader node (MODEL, CLIP, vocab_pack) → (MODEL, CLIP) from ComfyUI-Anima_lora-Adapter. Put the .safetensors and .json of a pack side by side and select the .safetensors. A regular LoRA loader cannot load it.
The current pack is anima_cjk_vocab_pack_preview5.safetensors plus its .json sidecar (per-glyph routing on; ComfyUI needs ComfyUI-Anima_lora-Adapter ≥ 3.13.0). It is preview4 with the component its trained rows share scaled to 0.8. 1,362 of its rows are trained to draw quoted Japanese text: 174 kana and punctuation, 1,185 kanji and 3 marks (゙ 々 〇), listed in anima_cjk_vocab_pack_preview5_trained.json. The earlier packs (preview … preview4) stay in the repo.
Intended use
Japanese in a prompt is text to draw in the image. Put the string you want rendered in double quotes inside an English clause, and keep the rest of the prompt in English tags:
1girl, solo, speech bubble. Text reads as "何言ってるの?".
1girl, solo, speech bubble, japanese text. Japanese text reads as "はい".
The japanese text tag (and the word Japanese in the clause) is optional. The Japanese rows are trained on rendering only, so Japanese written as tags (猫耳, 銀髪) is not the intended use: write those in English (cat ears, silver hair). Korean and Chinese spans map to tag rows; rendering them is not trained yet.
Why extend T5?
Anima can already draw many CJK glyphs. That ability lives in the DiT, but no prompt can reach it: the T5 tokenizer has no ids for Japanese, so a quoted "はい" never becomes a usable address.
This pack adds new T5-side token rows instead of changing the model. The rows are used only when a prompt contains CJK characters or symbols T5 cannot spell. Everything else tokenizes exactly as before, so English prompts are bit-identical with or without the pack.
The DiT, the LLM adapter and both text encoders stay untouched. Anything that works on base Anima still works, including LoRA training. For example, you can train a Japanese illustration LoRA with OCR text in the captions and use it together with the pack.
How it was trained
The new rows are the only trainable parameters. They are trained with the ordinary flow-matching loss, so the gradient flows back through the frozen DiT and the frozen adapter and lands only on the rows. The training images are scene composites with Japanese text in speech bubbles, and the loss is weighted up inside the text box.
Anima's text side, with every weight frozen. Qwen3 reads the raw caption and supplies the adapter's keys and values. The T5 id stream is the query: English keeps its stock ids, and each Japanese piece is routed to a pack row with id ≥ 32128. English captions never touch a pack row.
Curious how it works? See the write-up: Waking Anima up to read Japanese.
Usage in detail
This is a vocab pack, not a LoRA. It holds 69,558 extra text-embedding rows (ext_embed [69558, 1024], fp32) appended after the base table, so their ids start at 32128. The .json sidecar carries the tokenization and segmentation maps and the routing rule. Both files must sit together with the same stem.
A pack changes exactly two things in an Anima stack:
- The T5-side id stream. CJK spans are re-tokenized with the Qwen3 tokenizer and mapped to pack rows.
- The adapter's id table (
llm_adapter.embed). The pack rows are appended after the 32,128 stock rows.
The ComfyUI node and anima_lora do both steps for you. diffusers has no pack support yet, so there both steps are done by hand. The routing logic is one file, library/anima/ext_vocab.py, which needs only the standard library and torch. Install anima_lora or copy that file.
Runnable versions of both snippets are examples/09_cjk_vocab_pack.py (anima_lora engine) and examples/10_cjk_vocab_pack_diffusers.py (diffusers). Each has a --dry_run flag that prints the routed id stream without loading model weights.
diffusers (≥ 0.39, ModularPipeline)
import torch
from huggingface_hub import hf_hub_download
from diffusers.modular_pipelines.anima import AnimaAutoBlocks
from diffusers.modular_pipelines.anima.encoders import AnimaTextEncoderStep
from library.anima.ext_vocab import HybridT5Encoder, load_ext_assets # anima_lora repo
REPO, PACK = "sorryhyun/anima-vocab-pack-cjk", "anima_cjk_vocab_pack_preview5"
prefix = hf_hub_download(REPO, f"{PACK}.safetensors")[:-len(".safetensors")]
hf_hub_download(REPO, f"{PACK}.json")
table, mapping = load_ext_assets(prefix)
# 1. text_encoder block: T5 ids through the pack (stock ids for prompts with no routed character)
class PackTextEncoderStep(AnimaTextEncoderStep):
_enc = None
@classmethod
def _get_t5_prompt_ids(cls, components, prompt, max_sequence_length, device):
if cls._enc is None: # the pipeline's own tokenizers: same vocabularies the pack was built on
cls._enc = HybridT5Encoder.from_mapping(components.t5_tokenizer, components.tokenizer, mapping)
prompt = [prompt] if isinstance(prompt, str) else prompt
rows = []
for t in prompt:
if cls._enc.routes(t):
ids, mask = cls._enc.encode(t, max_sequence_length)
rows.append(ids[: sum(mask)])
else:
rows.append(components.t5_tokenizer(t, max_length=max_sequence_length, truncation=True)["input_ids"])
n = max(map(len, rows)); pad = components.t5_tokenizer.pad_token_id
ids = torch.tensor([r + [pad] * (n - len(r)) for r in rows])
mask = torch.tensor([[1] * len(r) + [0] * (n - len(r)) for r in rows])
return ids.to(device), mask.to(device)
blocks = AnimaAutoBlocks()
blocks.sub_blocks["text_encoder"] = PackTextEncoderStep()
pipe = blocks.init_pipeline("circlestone-labs/Anima-Base-v1.0-Diffusers")
pipe.load_components(torch_dtype=torch.bfloat16)
pipe.to("cuda")
# 2. widen the conditioner's id table once (dtype/device follow it)
emb = pipe.text_conditioner.embed
pipe.text_conditioner.embed = torch.nn.Embedding.from_pretrained(
torch.cat([emb.weight.data, table.to(emb.weight.dtype).to(emb.weight.device)])
)
image = pipe(prompt='1girl, solo, speech bubble. Text reads as "何言ってるの?".', num_inference_steps=30).images[0]
LoRAs loaded with pipe.load_lora_weights(...) compose with the pack, because they touch different parameters.
anima_lora
In anima_lora, a pack is one field on the request. The engine installs the routing tokenizer and appends the rows to llm_adapter.embed for you.
import torch
from huggingface_hub import hf_hub_download
from anima_lora import GenerationRequest, default_checkpoints, generate, get_generation_settings, load_vae, save_output
REPO, PACK = "sorryhyun/anima-vocab-pack-cjk", "anima_cjk_vocab_pack_preview5"
pack = hf_hub_download(REPO, f"{PACK}.safetensors")
hf_hub_download(REPO, f"{PACK}.json")
ckpt = default_checkpoints()
args = GenerationRequest(dit=ckpt.dit, vae=ckpt.vae, text_encoder=ckpt.text_encoder,
vocab_pack=pack, # a .safetensors path or its prefix; no_vocab_pack=True turns it off
prompt='1girl, solo, speech bubble. Text reads as "何言ってるの?".', save_path="out.png").to_args()
args.device = device = torch.device("cuda")
latent = generate(args, get_generation_settings(args))
vae = load_vae(args.vae, device="cpu", dtype=torch.bfloat16, eval=True)
save_output(args, vae, latent, device)
Training uses the same key: train.py --vocab_pack <prefix>. The LoRA is stamped with the pack digest (ss_ext_pack_sha), so loaders can warn when it is used with a different pack.
Limitations
- Rendering is at an early stage. As of 2026-09-30 (
preview4), single kana and kanji and short words render some of the time: on our five-string acceptance set (はい,おしい,やったネ,ちょっと来い,こんにちは, 16 renders each), 37 of 80 read back exactly, up from 8 before the retrain.preview5has not been measured on this set yet. Longer words and sentences are still unreliable. Results depend on the seed, so generate a few. - The pack cannot exceed the base model. It only gives prompts a way to reach glyphs Anima already knows. It cannot draw text more accurately or more cleanly than Anima itself can.
- Japanese tags may drift. The rendering rows were trained inside English caption frames such as
Text reads as "…". Japanese used in other ways, like tags, may behave differently than it did before rendering training. English prompts are unaffected. - Only
anima-base-v1.0was used for training. Other Anima checkpoints are untested.
Roadmap
- Reliable rendering of words and sentences
- Rendering for all 2,136 jōyō kanji
- Rendering for Korean and Chinese
Model tree for sorryhyun/anima-vocab-pack-cjk
Base model
nvidia/Cosmos-Predict2-2B-Text2Image