Lithuanian complete-annotation tagger: UPOS + full FEATS + lemmas

The complete-annotation member of the family: full UD morphological features (Definite, Polarity, PronType, … — not just agreement keys) and a real lemmatizer (edit-script head trained on 2.1M gold MATAS lemmas). Fine-tuned from EMBEDDIA/litlat-bert.

Which model do I want?

model best for emits
…-morphology highest UPOS/slots accuracy UPOS + core FEATS
this one complete annotations incl. lemmas full UD FEATS + lemmas
…-morphology-vdu accentuation pipelines traditional-grammar categories

Benchmarks (measured)

Official CoNLL-18 evaluation, gold tokenization, UD_Lithuanian-ALKSNIS test:

metric (F1) this model UDPipe 2 reference
UPOS 93.2 95.2
UFeats 84.1 89.1
Lemmas 94.7 92.9

The remaining UFeats gap vs UDPipe is dominated by peripheral-key conventions not derivable from the training sources (Aspect, Definite on numerals, Hyph). Same-protocol agreement-features projection: UPOS 96.7 / UFeats 89.1. Speed: 921 tok/s (ONNX INT8, single CPU).

Usage (verified) — note: custom two-head architecture

⚠️ AutoModelForTokenClassification.from_pretrained will not load this checkpoint correctly (it has a second, lemma-script head). Use the bundled ONNX model:

import json, numpy as np, onnxruntime as ort
from huggingface_hub import snapshot_download
from transformers import AutoTokenizer

d = snapshot_download("alexbalandi/litlat-bert-lithuanian-morphology-full")
tok = AutoTokenizer.from_pretrained(d)
sess = ort.InferenceSession(f"{d}/onnx/model_quantized.onnx",
                            providers=["CPUExecutionProvider"])
cfg = json.load(open(f"{d}/head_config.json", encoding="utf-8"))
labels, scripts = cfg["labels"], cfg["lemma_scripts"]

def apply_script(form, script):          # decode the lemma edit-script
    if script.startswith("W|"): return script.split("|", 1)[1]
    _, lower_first, strip, suffix = script.split("|", 3)
    base = (form[0].lower() + form[1:]) if int(lower_first) and form else form
    stem = base[:-int(strip)] if int(strip) else base
    return stem + suffix

words = ["Vilniuje", "gyvena", "graži", "mano", "sesuo"]
enc = tok(words, is_split_into_words=True, return_tensors="np")
label_logits, lemma_logits = sess.run(
    None, {k: v for k, v in enc.items() if k in ("input_ids", "attention_mask")})
seen = set()
for pos, wid in enumerate(enc.word_ids(0)):
    if wid is None or wid in seen: continue
    seen.add(wid)
    print(words[wid],
          apply_script(words[wid], scripts[int(lemma_logits[0, pos].argmax())]),
          labels[int(label_logits[0, pos].argmax())])
# Vilniuje Vilnius PROPN|Case=Loc|Gender=Masc|Number=Sing
# gyvena   gyventi VERB|Mood=Ind|Number=Sing|Person=3|Polarity=Pos|Tense=Pres|VerbForm=Fin
# mano     aš      PRON|Case=Gen|Definite=Ind|Number=Sing|Person=1|PronType=Prs
# ...

pytorch_model.bin holds the custom-wrapper state dict for continued training with the project's tooling (local/tagger-hf/head_modeling.py::load_custom_model).

Training data & lineage

Gold corpora only, no NC-licensed model anywhere in the lineage: MATAS v3.0 (CC BY 4.0) with full UD features reconstructed from its MULTEXT-East codes and Jablonskis XPOS, + UD_Lithuanian-ALKSNIS (CC BY-SA 4.0); contextual gaps filled by iterated constrained-decoding self-training; gold lemmas throughout for the lemma head.

License & attribution

CC BY-SA 4.0 (inherited from litlat-bert and ALKSNIS). Please credit: MATAS v3.0 (Rimkutė, Bielinskienė, Boizou, Dadurkevičius, Kovalevskaitė, Utka; CLARIN-LT), UD_Lithuanian-ALKSNIS (VDU), litlat-bert (Ulčar & Robnik-Šikonja, EMBEDDIA).

Downloads last month
17
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for alexbalandi/litlat-bert-lithuanian-morphology-full

Quantized
(4)
this model