Laya Multilingual for LiteRT β Android GPU and NPU
Run the multilingual checkpoint of convaiinnovations/laya
on a phone GPU or NPU with LiteRT. Laya reads a text or a JSON state and answers questions you define
at request time: pick one of several options, score on an ordinal scale, or give a
yes/no probability. Each question is one forward pass. Validated with LiteRT 2.2.0 on a
Samsung Galaxy S26 (SM-S942Q, Android 16): 51 ms of graph time per question at 256
tokens on the GPU with explicit FP32 computation, 36 ms on the Hexagon NPU, and 60 ms
end to end in the sample app on the GPU. On 201 English and Japanese question rows
the answers match the official laya 0.3.4 fp32 CPU implementation: the same argmax
on every choice and score question, and a maximum probability difference of 0.0014
on the GPU and 0.0069 on the NPU. Other Android GPUs and NPUs have not been validated
here.
Which of the three Laya packages to take. All three were tested on a Galaxy S26 GPU with LiteRT 2.2.0 at explicit FP32; the times below are graph time per question at a 256-token window, and the first two packages add the app's table lookup (about 15β20 ms):
- Laya-Multilingual-LiteRT: the multilingual checkpoint (mmBERT-base; the publisher lists 100+ languages; validated in English and Japanese). 0.68 GB on the phone, 51 ms per question (36 ms on the NPU). Has an Android sample app with a Kotlin tokenizer (
android/; also zero_shot_classification in litert-samples). Take this one for non-English or mixed text; it answers English too, close to the English checkpoint (Convai's XNLI English: 0.843 against 0.860). - Laya-English-LiteRT: the English checkpoint (ModernBERT-large) and its
typed-decisions/fine-tune for four workflows (customer service, invoice processing, security incidents, agent-trace observability). 0.85 GB each, 123 ms per question (126 ms for the fine-tune). No app of its own; the Kotlin GPU function on its card takes token ids from the app's own tokenizer. Take this one for English-only text when the higher English score matters, and its fine-tune for those four workflows. - laya-LiteRT: the same two checkpoints (not the fine-tune) in the token-id form: token ids in, embedding table inside the graph (the two packages above take embedding rows that the app looks up from a table file). Its GPU files keep fp32 weights: 1.3 GB (multilingual) or 1.7 GB (English), 54 ms (multilingual) or 127 ms (English) per question with no lookup to add. No app of its own; its card links an Android SDK with Kotlin tokenizers for both checkpoints. Take this one to feed token ids and skip the table lookup in the app.
This repository is the multilingual package.
Both panels are screenshots of the Android sample in android/, taken on the
Galaxy S26 with the GPU selected. The email is invented. The probabilities are the
values the model returned. assets/demo.mp4 shows the same run.
Files and supported configuration
| File | Window N | Bytes | Role |
|---|---|---|---|
laya_ml_s256_embeds_wfp16.tflite |
256 | 250,889,040 | recommended; validated on the S26 GPU, NPU and CPU |
laya_ml_s256_embeds_fp32.tflite |
256 | 500,970,372 | fp32 reference; validated on the S26 GPU |
laya_ml_s512_embeds_wfp16.tflite |
512 | 251,806,544 | longer texts; validated on desktop CPU only |
laya_ml_s512_embeds_fp32.tflite |
512 | 501,887,876 | fp32 reference; desktop CPU only |
laya_ml_act_head_fp32.tflite |
any | 795,816 | act head, shared by every main graph |
token_embeddings_fp16.bin, token_embeddings.json |
393,216,000 | [256000,768] float16 token table for the host lookup |
|
tokenizer.json, tokenizer_config.json |
34,363,188 | exact files from the pinned checkpoint | |
laya_ml_calibration.json |
9,156 | temperatures per question type and option count | |
laya_host.py |
Python host: prompt builder, embedding lookup, decoder | ||
android/ |
Android sample: Kotlin host and Compose UI | ||
fixtures/gate_rows_s256.json |
1,018,117 | the 201 validation rows with the reference answers |
The recommended phone set is the S256 wfp16 graph, the act head, the token table, the tokenizer pair and the calibration file: 679,274,525 bytes.
wfp16 files store the 99 FULLY_CONNECTED weight tensors as float16 with a
DEQUANTIZE to float32. Every activation and every other constant stays float32.
The graphs are safe to compute in fp16 (revision of 2026-09-29). The NPU computes in fp16, and three rewrites keep that finite. Each gives identical results in fp32: the graphs before and after agree bit for bit in PyTorch and give the same answers on every validation row. From layer 12 the start-of-text and separator rows carry values up to 1.4e4, whose squares overflow fp16 inside LayerNorm, so each large LayerNorm computes on its input scaled by a power of two with epsilon scaled by the square. The attention masks use β1e4 instead of β1e9, which is βinf in fp16 and turns a zero mask weight into NaN. Layers 11 and 12 move a power of two from the GeGLU product into the output projection. Without these rewrites every row was non-finite on the NPU and on the GPU at FP16.
The main graph holds the mmBERT-base encoder, Laya's two typed head layers and the
option scorer, applied at every position. The host does the rest: tokenize, build
the prompt with one <mask> marker per option, look up the embedding rows, read the
logits at the marker positions, and apply softmax. The act head is a separate small
graph because its input depends on those host-side probabilities.
The graphs take embedding rows, not token ids. With the token table inside the
graph, the float16-weight file does not compile on LiteRT 2.2.0 CompiledModel GPU: the
GPU delegate rejects the EMBEDDING_LOOKUP that reads a DEQUANTIZE-fed table
(Empty quantization params), and CompiledModel needs every operator on the GPU. The
fp32 file with the table inside does compile (1.29 GB, about 7 s of compilation;
litert-community/laya-LiteRT
ships that form). With the lookup on the host, the wfp16 graph compiles as one GPU
partition (1809 of 1809 operators) at 251 MB and 1.3 s. The float16 table reproduces
the checkpoint's float32 table exactly (maximum difference 0 over all 196,608,000
values).
On the GPU, use explicit FP32 computation, as in the Kotlin block below. GPU fp16
was measured and is not recommended: with FP32 accumulation one validation row missed
the 0.01 probability limit, and plain FP16 changed one of 81 answers. On the NPU, the
same file compiles on the phone through LiteRT's Qualcomm plugin (JIT): about 40 s on
the first launch, then about 0.2 s from the cache. The app must package the Qualcomm
runtime libraries; android/README.md lists them. INT8 files were not built. Questions
with more than 20 options are outside what was validated.
Minimal usage
Python β complete pipeline, desktop CPU
Needs numpy, transformers (tokenizer only) and ai-edge-litert. Neither torch
nor the laya package is required. Put laya_host.py next to the downloaded files.
import json
from pathlib import Path
from laya_host import LayaHost
assets = Path(".")
questions = {
"intent": {
"type": "choice",
"instructions": "What does the customer need?",
"criteria": {"refund": "money returned", "help": "technical help"},
}
}
with LayaHost(
tokenizer_dir=assets,
main_graph_path=assets / "laya_ml_s256_embeds_wfp16.tflite",
act_graph_path=assets / "laya_ml_act_head_fp32.tflite",
window=256,
head_max_len=256,
temperatures_json=assets / "laya_ml_calibration.json",
embeddings_path=assets / "token_embeddings_fp16.bin",
) as host:
result = host.predict("εγζ―ζγγδΊιγ«θ«ζ±γγγΎγγγθΏιγγι‘γγγΎγγ", questions)
print(json.dumps(result, ensure_ascii=False, indent=2))
assets is the directory with the downloaded files; tokenizer_dir needs
tokenizer.json and tokenizer_config.json. Run from a directory that held only the
seven files it names, this exact code printed:
"choice": "refund", "probabilities": {"refund": 0.9954, "help": 0.0046},
"confidence": 0.958, "action": {"act_probability": 1.0}, "usage": {"input_tokens": 37, ...}
Kotlin β Android GPU with explicit FP32
This is LayaMinimalUsage.kt from the sample, compiled into the app (one comment
trimmed). It gathers the float16 rows, writes the float32 buffers, runs both graphs
and decodes. LayaEngine
loads the tokenizer, the calibration file and the memory-mapped token table.
Build and install steps are in android/README.md.
import android.content.Context
import androidx.annotation.WorkerThread
object LayaMinimalUsage {
/** Call off the main thread after installing the files listed in README.md. */
@WorkerThread
fun classify(context: Context): Map<String, Any?> {
LayaEngine(context, LayaEngine.Storage.WFP16).use { engine ->
val question =
linkedMapOf<String, Any?>(
"type" to "choice",
"instructions" to "What does the customer need?",
"criteria" to linkedMapOf("refund" to "money returned", "help" to "technical help"),
)
val backend = LayaEngine.Backend.GPU
// CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32).
engine.initialize(backend)
val row = engine.prepare("εγζ―ζγγδΊιγ«θ«ζ±γγγΎγγγθΏιγγι‘γγγΎγγ", question)
val raw = engine.runRaw(row, backend)
check(raw.finite)
return LayaDecoder.decode(raw.markerLogits, raw.actLogits, row.question, engine.calibration)
}
}
}
For the NPU, use LayaEngine.Backend.NPU: LayaEngine then creates the environment with
DispatchLibraryDir and CompilerPluginLibraryDir set to the app's native library
directory, requests QualcommOptions(htpPerformanceMode = BURST) for the main graph, and
keeps the act head on the CPU.
Host contract
HOST_CONTRACT.md is the complete specification. In short:
- Build the sequence exactly as the upstream
build_sequencedoes:<bos> "<type> question: <instructions>" <eos> <mask> option0 <mask> option1 β¦ <eos> text <eos>. Each option keeps at most 48 tokens. The text is truncated on the right so the whole sequence fits in N. Remember the position of every<mask>. - Pad the ids on the right with id 0 to N.
inputs_embeds [1,N,768]holds the token-table row of every position, padding included.attention_mask [1,N]is 1 for real tokens.qtype_onehot [1,3]selects choice, score or yes/no. - The main graph returns
token_logits [1,N]andpooled_cls [1,768]. Take the logits at the marker positions, divide by the temperature for that question type and option count, and apply softmax. That is the answer distribution. - The act head takes
pooled_clsand four features of the uncalibrated distribution (top probability, top-two margin, normalized entropy, option count / 255) and returnsact_logits [1,2].
The sample app ships three upstream question presets unchanged: email triage, support intent and moderation.
Measured quality and performance
Reference: the official laya 0.3.4 predict on fp32 CPU at checkpoint revision
1c5edc17, with max_len set to the window. The inputs are 44 invented states
(21 English, 21 Japanese, 2 mixed) with the upstream presets, the upstream
quickstart schema and custom schemas of up to 20 options: 201 question rows, of which 81 are choice or score questions and 120 are
yes/no. Probabilities are compared at temperature 1, including the yes/no
probability and the act probability.
| Device, accelerator | Graph | Same argmax | Max probability difference | Placement | Median ms per question |
|---|---|---|---|---|---|
| S26 GPU FP32 | S256 wfp16 | 81/81 | 0.0014 | 1809/1809, one partition | 51.9 |
| S26 NPU (Hexagon v81, JIT) | S256 wfp16 | 81/81 | 0.0069 | whole graph, one NPU node | 36.5 |
| S26 GPU FP32, previous revision | S256 fp32 | 81/81 | 0.0001 | 1680/1680, one partition | 50.8 |
| S26 CPU, previous revision | S256 wfp16 | 81/81 | 0.0014 | 163.0 | |
| Desktop CPU | S256 and S512, wfp16 | 81/81 | 0.0014 | ||
| Desktop CPU | S256 and S512, fp32 | 81/81 | 0.0001 |
On the phone the Kotlin tokenizer and prompt builder produced the same token ids and marker positions as the Python reference on all 201 rows. The previous-revision rows were measured on 2026-09-21 with the graphs before the fp16 rewrites, which give the same fp32 results. Desktop CPU is ai-edge-litert 2.1.6 on macOS arm64 with 4 threads.
The per-question time covers both graphs, from writing the inputs to the end of
output readback: the median of 200 rows after one cold call, debug build, screen on,
no other workload, battery at 35β37 Β°C. The host embedding lookup adds a median of
14.6 ms. In the non-debuggable build of the sample, a five-question email-triage run
took 301 ms and 303 ms in two fresh launches, and the app reached Ready 2.3 s after
onCreate. That start-up is 1.26 s of GPU compilation, 0.76 s of tokenizer loading
and a 0.27 s warm-up pass. With the NPU selected, the debug build answered the same run
in 257 ms and 268 ms and reached Ready 2.0 s after onCreate, of which 0.23 s loaded
the compiled NPU graph from the cache. These are samples from one device, not a
benchmark across devices or thermal states.
Calibration
The upstream multilingual checkpoint ships with every temperature at 1.
laya_ml_calibration.json holds temperatures fitted here, one per question type and
option-count bucket, on 4,415 labeled English and Japanese examples from public
datasets, at the 256-token window. A temperature never changes which option wins.
| Bucket | Validation rows | Temperature | ECE at T=1 β shipped |
|---|---|---|---|
| choice, 2 options | 140 | 1.40 | 0.089 β 0.117 |
| choice, 3β5 | 150 | 1.36 | 0.154 β 0.087 |
| choice, 6β10 | 125 | 1.00 | 0.162 β 0.162 |
| choice, 11β20 | 108 | 2.45 | 0.383 β 0.120 |
| score, 3β5 levels | 161 | 3.68 | 0.257 β 0.087 |
| yes/no | 200 | 3.63 | 0.310 β 0.200 |
Fits use the same number of English and Japanese rows. A bucket keeps T=1 when a fitted value makes either language worse by more than 0.05 ECE. That is the case for 6β10 options. For 2 options the fit helps Japanese and costs English 0.03, so the pooled figure rises. This is a generic starting point: refit on your own data before you rely on the probabilities. Sources, licenses and label mappings are in licenses/CALIBRATION_DATASETS.md.
Limits
- The numbers above measure agreement with the upstream model, not task accuracy.
The conversion reproduces the upstream answers, including the wrong ones. On five
of the Japanese rows the upstream answer differs from the intended reading; one
example is a veiled threat that scores 0.0 for
threat. They are listed inHOST_CONTRACT.md. - The act probability was 1.0 on all 201 rows, in the upstream model and in these graphs. Do not treat it as an escalation signal without your own evaluation.
- The S512 graphs were checked on desktop CPU only. The calibration was fitted at 256 tokens.
- One device was tested: one Snapdragon 8 Elite Gen 5 phone for the GPU and the NPU. The token table is memory-mapped and the app needs about 680 MB of files in its private storage. The NPU libraries add about 40 MB to the APK.
Provenance, conversion and license
- Source:
convaiinnovations/laya, revision1c5edc17a7acd8701df6fc341c0d179f1c62c982, subfoldermultilingual/(encoder jhu-clsp/mmBERT-base, 22 layers, hidden 768, vocabulary 256,000), loaded withlaya0.3.4 andtransformers5.17.0. - Conversion:
litert-torch0.9.3 (torch2.12.1), fixed shapes, fp32. The math is unchanged: host-side embedding lookup, the question type as a one-hot matmul, padding and sliding-window masks as float constants, rotary tables baked per layer type, attention kept at rank 4, the two head layers written out explicitly, exact GELU, and the scorer applied at every position. The 2026-09-29 revision adds the three fp16 rewrites described above, all exact in fp32. - Weight storage:
ai-edge-quantizer0.8.0 float16 FLOAT_CASTING on the FULLY_CONNECTED weights only (wfp16files). - Verification: LiteRT CompiledModel Python API on desktop CPU, and the CompiledModel Kotlin API on the Galaxy S26 for CPU, GPU and NPU. Every gate is the same argmax plus an absolute probability difference. Correlation was never used as a gate.
License: the Laya checkpoint and code are Apache-2.0, and these converted files are
released under the same license. mmBERT-base declares MIT. The host code ports logic
from laya, transformers and tokenizers (Apache-2.0). License texts are in
licenses/, and attribution is in NOTICE. The calibration datasets keep their own
terms; none of their text is included here.
- Downloads last month
- 423
Model tree for litert-community/Laya-Multilingual-LiteRT
Base model
convaiinnovations/laya