Laya Multilingual for LiteRT β€” Android GPU and NPU

Run the multilingual checkpoint of convaiinnovations/laya on a phone GPU or NPU with LiteRT. Laya reads a text or a JSON state and answers questions you define at request time: pick one of several options, score on an ordinal scale, or give a yes/no probability. Each question is one forward pass. Validated with LiteRT 2.2.0 on a Samsung Galaxy S26 (SM-S942Q, Android 16): 51 ms of graph time per question at 256 tokens on the GPU with explicit FP32 computation, 36 ms on the Hexagon NPU, and 60 ms end to end in the sample app on the GPU. On 201 English and Japanese question rows the answers match the official laya 0.3.4 fp32 CPU implementation: the same argmax on every choice and score question, and a maximum probability difference of 0.0014 on the GPU and 0.0069 on the NPU. Other Android GPUs and NPUs have not been validated here.

Which of the three Laya packages to take. All three were tested on a Galaxy S26 GPU with LiteRT 2.2.0 at explicit FP32; the times below are graph time per question at a 256-token window, and the first two packages add the app's table lookup (about 15–20 ms):

  • Laya-Multilingual-LiteRT: the multilingual checkpoint (mmBERT-base; the publisher lists 100+ languages; validated in English and Japanese). 0.68 GB on the phone, 51 ms per question (36 ms on the NPU). Has an Android sample app with a Kotlin tokenizer (android/; also zero_shot_classification in litert-samples). Take this one for non-English or mixed text; it answers English too, close to the English checkpoint (Convai's XNLI English: 0.843 against 0.860).
  • Laya-English-LiteRT: the English checkpoint (ModernBERT-large) and its typed-decisions/ fine-tune for four workflows (customer service, invoice processing, security incidents, agent-trace observability). 0.85 GB each, 123 ms per question (126 ms for the fine-tune). No app of its own; the Kotlin GPU function on its card takes token ids from the app's own tokenizer. Take this one for English-only text when the higher English score matters, and its fine-tune for those four workflows.
  • laya-LiteRT: the same two checkpoints (not the fine-tune) in the token-id form: token ids in, embedding table inside the graph (the two packages above take embedding rows that the app looks up from a table file). Its GPU files keep fp32 weights: 1.3 GB (multilingual) or 1.7 GB (English), 54 ms (multilingual) or 127 ms (English) per question with no lookup to add. No app of its own; its card links an Android SDK with Kotlin tokenizers for both checkpoints. Take this one to feed token ids and skip the table lookup in the app.

This repository is the multilingual package.

The sample app: an invented Japanese support email on the left, the answer to the first question on the right

Both panels are screenshots of the Android sample in android/, taken on the Galaxy S26 with the GPU selected. The email is invented. The probabilities are the values the model returned. assets/demo.mp4 shows the same run.

Files and supported configuration

File Window N Bytes Role
laya_ml_s256_embeds_wfp16.tflite 256 250,889,040 recommended; validated on the S26 GPU, NPU and CPU
laya_ml_s256_embeds_fp32.tflite 256 500,970,372 fp32 reference; validated on the S26 GPU
laya_ml_s512_embeds_wfp16.tflite 512 251,806,544 longer texts; validated on desktop CPU only
laya_ml_s512_embeds_fp32.tflite 512 501,887,876 fp32 reference; desktop CPU only
laya_ml_act_head_fp32.tflite any 795,816 act head, shared by every main graph
token_embeddings_fp16.bin, token_embeddings.json 393,216,000 [256000,768] float16 token table for the host lookup
tokenizer.json, tokenizer_config.json 34,363,188 exact files from the pinned checkpoint
laya_ml_calibration.json 9,156 temperatures per question type and option count
laya_host.py Python host: prompt builder, embedding lookup, decoder
android/ Android sample: Kotlin host and Compose UI
fixtures/gate_rows_s256.json 1,018,117 the 201 validation rows with the reference answers

The recommended phone set is the S256 wfp16 graph, the act head, the token table, the tokenizer pair and the calibration file: 679,274,525 bytes.

wfp16 files store the 99 FULLY_CONNECTED weight tensors as float16 with a DEQUANTIZE to float32. Every activation and every other constant stays float32.

The graphs are safe to compute in fp16 (revision of 2026-09-29). The NPU computes in fp16, and three rewrites keep that finite. Each gives identical results in fp32: the graphs before and after agree bit for bit in PyTorch and give the same answers on every validation row. From layer 12 the start-of-text and separator rows carry values up to 1.4e4, whose squares overflow fp16 inside LayerNorm, so each large LayerNorm computes on its input scaled by a power of two with epsilon scaled by the square. The attention masks use βˆ’1e4 instead of βˆ’1e9, which is βˆ’inf in fp16 and turns a zero mask weight into NaN. Layers 11 and 12 move a power of two from the GeGLU product into the output projection. Without these rewrites every row was non-finite on the NPU and on the GPU at FP16.

The main graph holds the mmBERT-base encoder, Laya's two typed head layers and the option scorer, applied at every position. The host does the rest: tokenize, build the prompt with one <mask> marker per option, look up the embedding rows, read the logits at the marker positions, and apply softmax. The act head is a separate small graph because its input depends on those host-side probabilities.

The graphs take embedding rows, not token ids. With the token table inside the graph, the float16-weight file does not compile on LiteRT 2.2.0 CompiledModel GPU: the GPU delegate rejects the EMBEDDING_LOOKUP that reads a DEQUANTIZE-fed table (Empty quantization params), and CompiledModel needs every operator on the GPU. The fp32 file with the table inside does compile (1.29 GB, about 7 s of compilation; litert-community/laya-LiteRT ships that form). With the lookup on the host, the wfp16 graph compiles as one GPU partition (1809 of 1809 operators) at 251 MB and 1.3 s. The float16 table reproduces the checkpoint's float32 table exactly (maximum difference 0 over all 196,608,000 values).

On the GPU, use explicit FP32 computation, as in the Kotlin block below. GPU fp16 was measured and is not recommended: with FP32 accumulation one validation row missed the 0.01 probability limit, and plain FP16 changed one of 81 answers. On the NPU, the same file compiles on the phone through LiteRT's Qualcomm plugin (JIT): about 40 s on the first launch, then about 0.2 s from the cache. The app must package the Qualcomm runtime libraries; android/README.md lists them. INT8 files were not built. Questions with more than 20 options are outside what was validated.

Minimal usage

Python β€” complete pipeline, desktop CPU

Needs numpy, transformers (tokenizer only) and ai-edge-litert. Neither torch nor the laya package is required. Put laya_host.py next to the downloaded files.

import json
from pathlib import Path
from laya_host import LayaHost

assets = Path(".")
questions = {
    "intent": {
        "type": "choice",
        "instructions": "What does the customer need?",
        "criteria": {"refund": "money returned", "help": "technical help"},
    }
}
with LayaHost(
    tokenizer_dir=assets,
    main_graph_path=assets / "laya_ml_s256_embeds_wfp16.tflite",
    act_graph_path=assets / "laya_ml_act_head_fp32.tflite",
    window=256,
    head_max_len=256,
    temperatures_json=assets / "laya_ml_calibration.json",
    embeddings_path=assets / "token_embeddings_fp16.bin",
) as host:
    result = host.predict("εŒγ˜ζ”―ζ‰•γ„γŒδΊŒι‡γ«θ«‹ζ±‚γ•γ‚ŒγΎγ—γŸγ€‚θΏ”ι‡‘γ‚’γŠι‘˜γ„γ—γΎγ™γ€‚", questions)
    print(json.dumps(result, ensure_ascii=False, indent=2))

assets is the directory with the downloaded files; tokenizer_dir needs tokenizer.json and tokenizer_config.json. Run from a directory that held only the seven files it names, this exact code printed:

"choice": "refund", "probabilities": {"refund": 0.9954, "help": 0.0046},
"confidence": 0.958, "action": {"act_probability": 1.0}, "usage": {"input_tokens": 37, ...}

Kotlin β€” Android GPU with explicit FP32

This is LayaMinimalUsage.kt from the sample, compiled into the app (one comment trimmed). It gathers the float16 rows, writes the float32 buffers, runs both graphs and decodes. LayaEngine loads the tokenizer, the calibration file and the memory-mapped token table. Build and install steps are in android/README.md.

import android.content.Context
import androidx.annotation.WorkerThread

object LayaMinimalUsage {
  /** Call off the main thread after installing the files listed in README.md. */
  @WorkerThread
  fun classify(context: Context): Map<String, Any?> {
    LayaEngine(context, LayaEngine.Storage.WFP16).use { engine ->
      val question =
        linkedMapOf<String, Any?>(
          "type" to "choice",
          "instructions" to "What does the customer need?",
          "criteria" to linkedMapOf("refund" to "money returned", "help" to "technical help"),
        )
      val backend = LayaEngine.Backend.GPU
      // CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32).
      engine.initialize(backend)
      val row = engine.prepare("εŒγ˜ζ”―ζ‰•γ„γŒδΊŒι‡γ«θ«‹ζ±‚γ•γ‚ŒγΎγ—γŸγ€‚θΏ”ι‡‘γ‚’γŠι‘˜γ„γ—γΎγ™γ€‚", question)
      val raw = engine.runRaw(row, backend)
      check(raw.finite)
      return LayaDecoder.decode(raw.markerLogits, raw.actLogits, row.question, engine.calibration)
    }
  }
}

For the NPU, use LayaEngine.Backend.NPU: LayaEngine then creates the environment with DispatchLibraryDir and CompilerPluginLibraryDir set to the app's native library directory, requests QualcommOptions(htpPerformanceMode = BURST) for the main graph, and keeps the act head on the CPU.

Host contract

HOST_CONTRACT.md is the complete specification. In short:

  1. Build the sequence exactly as the upstream build_sequence does: <bos> "<type> question: <instructions>" <eos> <mask> option0 <mask> option1 … <eos> text <eos>. Each option keeps at most 48 tokens. The text is truncated on the right so the whole sequence fits in N. Remember the position of every <mask>.
  2. Pad the ids on the right with id 0 to N. inputs_embeds [1,N,768] holds the token-table row of every position, padding included. attention_mask [1,N] is 1 for real tokens. qtype_onehot [1,3] selects choice, score or yes/no.
  3. The main graph returns token_logits [1,N] and pooled_cls [1,768]. Take the logits at the marker positions, divide by the temperature for that question type and option count, and apply softmax. That is the answer distribution.
  4. The act head takes pooled_cls and four features of the uncalibrated distribution (top probability, top-two margin, normalized entropy, option count / 255) and returns act_logits [1,2].

The sample app ships three upstream question presets unchanged: email triage, support intent and moderation.

Measured quality and performance

Reference: the official laya 0.3.4 predict on fp32 CPU at checkpoint revision 1c5edc17, with max_len set to the window. The inputs are 44 invented states (21 English, 21 Japanese, 2 mixed) with the upstream presets, the upstream quickstart schema and custom schemas of up to 20 options: 201 question rows, of which 81 are choice or score questions and 120 are yes/no. Probabilities are compared at temperature 1, including the yes/no probability and the act probability.

Device, accelerator Graph Same argmax Max probability difference Placement Median ms per question
S26 GPU FP32 S256 wfp16 81/81 0.0014 1809/1809, one partition 51.9
S26 NPU (Hexagon v81, JIT) S256 wfp16 81/81 0.0069 whole graph, one NPU node 36.5
S26 GPU FP32, previous revision S256 fp32 81/81 0.0001 1680/1680, one partition 50.8
S26 CPU, previous revision S256 wfp16 81/81 0.0014 163.0
Desktop CPU S256 and S512, wfp16 81/81 0.0014
Desktop CPU S256 and S512, fp32 81/81 0.0001

On the phone the Kotlin tokenizer and prompt builder produced the same token ids and marker positions as the Python reference on all 201 rows. The previous-revision rows were measured on 2026-09-21 with the graphs before the fp16 rewrites, which give the same fp32 results. Desktop CPU is ai-edge-litert 2.1.6 on macOS arm64 with 4 threads.

The per-question time covers both graphs, from writing the inputs to the end of output readback: the median of 200 rows after one cold call, debug build, screen on, no other workload, battery at 35–37 Β°C. The host embedding lookup adds a median of 14.6 ms. In the non-debuggable build of the sample, a five-question email-triage run took 301 ms and 303 ms in two fresh launches, and the app reached Ready 2.3 s after onCreate. That start-up is 1.26 s of GPU compilation, 0.76 s of tokenizer loading and a 0.27 s warm-up pass. With the NPU selected, the debug build answered the same run in 257 ms and 268 ms and reached Ready 2.0 s after onCreate, of which 0.23 s loaded the compiled NPU graph from the cache. These are samples from one device, not a benchmark across devices or thermal states.

Calibration

The upstream multilingual checkpoint ships with every temperature at 1. laya_ml_calibration.json holds temperatures fitted here, one per question type and option-count bucket, on 4,415 labeled English and Japanese examples from public datasets, at the 256-token window. A temperature never changes which option wins.

Bucket Validation rows Temperature ECE at T=1 β†’ shipped
choice, 2 options 140 1.40 0.089 β†’ 0.117
choice, 3–5 150 1.36 0.154 β†’ 0.087
choice, 6–10 125 1.00 0.162 β†’ 0.162
choice, 11–20 108 2.45 0.383 β†’ 0.120
score, 3–5 levels 161 3.68 0.257 β†’ 0.087
yes/no 200 3.63 0.310 β†’ 0.200

Fits use the same number of English and Japanese rows. A bucket keeps T=1 when a fitted value makes either language worse by more than 0.05 ECE. That is the case for 6–10 options. For 2 options the fit helps Japanese and costs English 0.03, so the pooled figure rises. This is a generic starting point: refit on your own data before you rely on the probabilities. Sources, licenses and label mappings are in licenses/CALIBRATION_DATASETS.md.

Limits

  • The numbers above measure agreement with the upstream model, not task accuracy. The conversion reproduces the upstream answers, including the wrong ones. On five of the Japanese rows the upstream answer differs from the intended reading; one example is a veiled threat that scores 0.0 for threat. They are listed in HOST_CONTRACT.md.
  • The act probability was 1.0 on all 201 rows, in the upstream model and in these graphs. Do not treat it as an escalation signal without your own evaluation.
  • The S512 graphs were checked on desktop CPU only. The calibration was fitted at 256 tokens.
  • One device was tested: one Snapdragon 8 Elite Gen 5 phone for the GPU and the NPU. The token table is memory-mapped and the app needs about 680 MB of files in its private storage. The NPU libraries add about 40 MB to the APK.

Provenance, conversion and license

  • Source: convaiinnovations/laya, revision 1c5edc17a7acd8701df6fc341c0d179f1c62c982, subfolder multilingual/ (encoder jhu-clsp/mmBERT-base, 22 layers, hidden 768, vocabulary 256,000), loaded with laya 0.3.4 and transformers 5.17.0.
  • Conversion: litert-torch 0.9.3 (torch 2.12.1), fixed shapes, fp32. The math is unchanged: host-side embedding lookup, the question type as a one-hot matmul, padding and sliding-window masks as float constants, rotary tables baked per layer type, attention kept at rank 4, the two head layers written out explicitly, exact GELU, and the scorer applied at every position. The 2026-09-29 revision adds the three fp16 rewrites described above, all exact in fp32.
  • Weight storage: ai-edge-quantizer 0.8.0 float16 FLOAT_CASTING on the FULLY_CONNECTED weights only (wfp16 files).
  • Verification: LiteRT CompiledModel Python API on desktop CPU, and the CompiledModel Kotlin API on the Galaxy S26 for CPU, GPU and NPU. Every gate is the same argmax plus an absolute probability difference. Correlation was never used as a gate.

License: the Laya checkpoint and code are Apache-2.0, and these converted files are released under the same license. mmBERT-base declares MIT. The host code ports logic from laya, transformers and tokenizers (Apache-2.0). License texts are in licenses/, and attribution is in NOTICE. The calibration datasets keep their own terms; none of their text is included here.

Downloads last month
423
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for litert-community/Laya-Multilingual-LiteRT

Finetuned
(101)
this model