Lithuanian complete-annotation tagger: UPOS + full FEATS + lemmas
The complete-annotation member of the family: full UD morphological features (Definite, Polarity, PronType, … — not just agreement keys) and a real lemmatizer (edit-script head trained on 2.1M gold MATAS lemmas). Fine-tuned from EMBEDDIA/litlat-bert.
Which model do I want?
| model | best for | emits |
|---|---|---|
| …-morphology | highest UPOS/slots accuracy | UPOS + core FEATS |
| this one | complete annotations incl. lemmas | full UD FEATS + lemmas |
| …-morphology-vdu | accentuation pipelines | traditional-grammar categories |
Benchmarks (measured)
Official CoNLL-18 evaluation, gold tokenization, UD_Lithuanian-ALKSNIS test:
| metric (F1) | this model | UDPipe 2 reference |
|---|---|---|
| UPOS | 93.2 | 95.2 |
| UFeats | 84.1 | 89.1 |
| Lemmas | 94.7 | 92.9 |
The remaining UFeats gap vs UDPipe is dominated by peripheral-key conventions not derivable from the training sources (Aspect, Definite on numerals, Hyph). Same-protocol agreement-features projection: UPOS 96.7 / UFeats 89.1. Speed: 921 tok/s (ONNX INT8, single CPU).
Usage (verified) — note: custom two-head architecture
⚠️ AutoModelForTokenClassification.from_pretrained will not load
this checkpoint correctly (it has a second, lemma-script head). Use the
bundled ONNX model:
import json, numpy as np, onnxruntime as ort
from huggingface_hub import snapshot_download
from transformers import AutoTokenizer
d = snapshot_download("alexbalandi/litlat-bert-lithuanian-morphology-full")
tok = AutoTokenizer.from_pretrained(d)
sess = ort.InferenceSession(f"{d}/onnx/model_quantized.onnx",
providers=["CPUExecutionProvider"])
cfg = json.load(open(f"{d}/head_config.json", encoding="utf-8"))
labels, scripts = cfg["labels"], cfg["lemma_scripts"]
def apply_script(form, script): # decode the lemma edit-script
if script.startswith("W|"): return script.split("|", 1)[1]
_, lower_first, strip, suffix = script.split("|", 3)
base = (form[0].lower() + form[1:]) if int(lower_first) and form else form
stem = base[:-int(strip)] if int(strip) else base
return stem + suffix
words = ["Vilniuje", "gyvena", "graži", "mano", "sesuo"]
enc = tok(words, is_split_into_words=True, return_tensors="np")
label_logits, lemma_logits = sess.run(
None, {k: v for k, v in enc.items() if k in ("input_ids", "attention_mask")})
seen = set()
for pos, wid in enumerate(enc.word_ids(0)):
if wid is None or wid in seen: continue
seen.add(wid)
print(words[wid],
apply_script(words[wid], scripts[int(lemma_logits[0, pos].argmax())]),
labels[int(label_logits[0, pos].argmax())])
# Vilniuje Vilnius PROPN|Case=Loc|Gender=Masc|Number=Sing
# gyvena gyventi VERB|Mood=Ind|Number=Sing|Person=3|Polarity=Pos|Tense=Pres|VerbForm=Fin
# mano aš PRON|Case=Gen|Definite=Ind|Number=Sing|Person=1|PronType=Prs
# ...
pytorch_model.bin holds the custom-wrapper state dict for continued
training with the project's tooling
(local/tagger-hf/head_modeling.py::load_custom_model).
Training data & lineage
Gold corpora only, no NC-licensed model anywhere in the lineage: MATAS v3.0 (CC BY 4.0) with full UD features reconstructed from its MULTEXT-East codes and Jablonskis XPOS, + UD_Lithuanian-ALKSNIS (CC BY-SA 4.0); contextual gaps filled by iterated constrained-decoding self-training; gold lemmas throughout for the lemma head.
License & attribution
CC BY-SA 4.0 (inherited from litlat-bert and ALKSNIS). Please credit: MATAS v3.0 (Rimkutė, Bielinskienė, Boizou, Dadurkevičius, Kovalevskaitė, Utka; CLARIN-LT), UD_Lithuanian-ALKSNIS (VDU), litlat-bert (Ulčar & Robnik-Šikonja, EMBEDDIA).
- Downloads last month
- 17
Model tree for alexbalandi/litlat-bert-lithuanian-morphology-full
Base model
EMBEDDIA/litlat-bert