--- license: apache-2.0 pipeline_tag: text-classification language: - en - ja base_model: - convaiinnovations/laya tags: - litert - android - text-classification - zero-shot-classification - laya - modernbert --- # Laya Multilingual for LiteRT — Android GPU and NPU Run the multilingual checkpoint of [convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya) on a phone GPU or NPU with [LiteRT](https://github.com/google-ai-edge/litert). Laya reads a text or a JSON state and answers questions you define at request time: pick one of several options, score on an ordinal scale, or give a yes/no probability. Each question is one forward pass. **Validated with LiteRT 2.2.0 on a Samsung Galaxy S26 (SM-S942Q, Android 16): 51 ms of graph time per question at 256 tokens on the GPU with explicit FP32 computation, 36 ms on the Hexagon NPU, and 60 ms end to end in the sample app on the GPU.** On 201 English and Japanese question rows the answers match the official `laya` 0.3.4 fp32 CPU implementation: the same argmax on every choice and score question, and a maximum probability difference of 0.0014 on the GPU and 0.0069 on the NPU. Other Android GPUs and NPUs have not been validated here. **Which of the three Laya packages to take.** All three were tested on a Galaxy S26 GPU with [LiteRT](https://github.com/google-ai-edge/litert) 2.2.0 at explicit FP32; the times below are graph time per question at a 256-token window, and the first two packages add the app's table lookup (about 15–20 ms): - [Laya-Multilingual-LiteRT](https://huggingface.co/litert-community/Laya-Multilingual-LiteRT): the multilingual checkpoint (mmBERT-base; the publisher lists 100+ languages; validated in English and Japanese). 0.68 GB on the phone, 51 ms per question (36 ms on the NPU). Has an Android sample app with a Kotlin tokenizer (`android/`; also [zero_shot_classification](https://github.com/google-ai-edge/litert-samples/tree/main/samples/litert/zero_shot_classification) in litert-samples). Take this one for non-English or mixed text; it answers English too, close to the English checkpoint (Convai's XNLI English: 0.843 against 0.860). - [Laya-English-LiteRT](https://huggingface.co/litert-community/Laya-English-LiteRT): the English checkpoint (ModernBERT-large) and its `typed-decisions/` fine-tune for four workflows (customer service, invoice processing, security incidents, agent-trace observability). 0.85 GB each, 123 ms per question (126 ms for the fine-tune). No app of its own; the Kotlin GPU function on its card takes token ids from the app's own tokenizer. Take this one for English-only text when the higher English score matters, and its fine-tune for those four workflows. - [laya-LiteRT](https://huggingface.co/litert-community/laya-LiteRT): the same two checkpoints (not the fine-tune) in the token-id form: token ids in, embedding table inside the graph (the two packages above take embedding rows that the app looks up from a table file). Its GPU files keep fp32 weights: 1.3 GB (multilingual) or 1.7 GB (English), 54 ms (multilingual) or 127 ms (English) per question with no lookup to add. No app of its own; it carries a reference host (`laya_host.py`) and `HOST_CONTRACT.md`, which specifies the sequence and the decoding for a port. Take this one to feed token ids and skip the table lookup in the app. This repository is the multilingual package. ![The sample app: an invented Japanese support email on the left, the answer to the first question on the right](assets/hero.png) Both panels are screenshots of the Android sample in `android/`, taken on the Galaxy S26 with the GPU selected. The email is invented. The probabilities are the values the model returned. [assets/demo.mp4](assets/demo.mp4) shows the same run. ## Files and supported configuration | File | Window N | Bytes | Role | |---|---:|---:|---| | `laya_ml_s256_embeds_wfp16.tflite` | 256 | 250,889,040 | recommended; validated on the S26 GPU, NPU and CPU | | `laya_ml_s256_embeds_fp32.tflite` | 256 | 500,970,372 | fp32 reference; validated on the S26 GPU | | `laya_ml_s512_embeds_wfp16.tflite` | 512 | 251,806,544 | longer texts; validated on desktop CPU only | | `laya_ml_s512_embeds_fp32.tflite` | 512 | 501,887,876 | fp32 reference; desktop CPU only | | `laya_ml_act_head_fp32.tflite` | any | 795,816 | act head, shared by every main graph | | `token_embeddings_fp16.bin`, `token_embeddings.json` | | 393,216,000 | `[256000,768]` float16 token table for the host lookup | | `tokenizer.json`, `tokenizer_config.json` | | 34,363,188 | exact files from the pinned checkpoint | | `laya_ml_calibration.json` | | 9,156 | temperatures per question type and option count | | `laya_host.py` | | | Python host: prompt builder, embedding lookup, decoder | | `android/` | | | Android sample: Kotlin host and Compose UI | | `fixtures/gate_rows_s256.json` | | 1,018,117 | the 201 validation rows with the reference answers | The recommended phone set is the S256 wfp16 graph, the act head, the token table, the tokenizer pair and the calibration file: 679,274,525 bytes. `wfp16` files store the 99 FULLY_CONNECTED weight tensors as float16 with a DEQUANTIZE to float32. Every activation and every other constant stays float32. **The graphs are safe to compute in fp16 (revision of 2026-09-29).** The NPU computes in fp16, and three rewrites keep that finite. Each gives identical results in fp32: the graphs before and after agree bit for bit in PyTorch and give the same answers on every validation row. From layer 12 the start-of-text and separator rows carry values up to 1.4e4, whose squares overflow fp16 inside LayerNorm, so each large LayerNorm computes on its input scaled by a power of two with epsilon scaled by the square. The attention masks use −1e4 instead of −1e9, which is −inf in fp16 and turns a zero mask weight into NaN. Layers 11 and 12 move a power of two from the GeGLU product into the output projection. Without these rewrites every row was non-finite on the NPU and on the GPU at FP16. The main graph holds the mmBERT-base encoder, Laya's two typed head layers and the option scorer, applied at every position. The host does the rest: tokenize, build the prompt with one `` marker per option, look up the embedding rows, read the logits at the marker positions, and apply softmax. The act head is a separate small graph because its input depends on those host-side probabilities. **The graphs take embedding rows, not token ids.** With the token table inside the graph, the float16-weight file does not compile on LiteRT 2.2.0 CompiledModel GPU: the GPU delegate rejects the EMBEDDING_LOOKUP that reads a DEQUANTIZE-fed table (`Empty quantization params`), and CompiledModel needs every operator on the GPU. The fp32 file with the table inside does compile (1.29 GB, about 7 s of compilation; [litert-community/laya-LiteRT](https://huggingface.co/litert-community/laya-LiteRT) ships that form). With the lookup on the host, the wfp16 graph compiles as one GPU partition (1809 of 1809 operators) at 251 MB and 1.3 s. The float16 table reproduces the checkpoint's float32 table exactly (maximum difference 0 over all 196,608,000 values). **On the GPU, use explicit FP32 computation**, as in the Kotlin block below. GPU fp16 was measured and is not recommended: with FP32 accumulation one validation row missed the 0.01 probability limit, and plain FP16 changed one of 81 answers. **On the NPU**, the same file compiles on the phone through LiteRT's Qualcomm plugin (JIT): about 40 s on the first launch, then about 0.2 s from the cache. The app must package the Qualcomm runtime libraries; `android/README.md` lists them. INT8 files were not built. Questions with more than 20 options are outside what was validated. ## Minimal usage ### Python — complete pipeline, desktop CPU Needs `numpy`, `transformers` (tokenizer only) and `ai-edge-litert`. Neither torch nor the `laya` package is required. Put `laya_host.py` next to the downloaded files. ```python import json from pathlib import Path from laya_host import LayaHost assets = Path(".") questions = { "intent": { "type": "choice", "instructions": "What does the customer need?", "criteria": {"refund": "money returned", "help": "technical help"}, } } with LayaHost( tokenizer_dir=assets, main_graph_path=assets / "laya_ml_s256_embeds_wfp16.tflite", act_graph_path=assets / "laya_ml_act_head_fp32.tflite", window=256, head_max_len=256, temperatures_json=assets / "laya_ml_calibration.json", embeddings_path=assets / "token_embeddings_fp16.bin", ) as host: result = host.predict("同じ支払いが二重に請求されました。返金をお願いします。", questions) print(json.dumps(result, ensure_ascii=False, indent=2)) ``` `assets` is the directory with the downloaded files; `tokenizer_dir` needs `tokenizer.json` and `tokenizer_config.json`. Run from a directory that held only the seven files it names, this exact code printed: ```text "choice": "refund", "probabilities": {"refund": 0.9954, "help": 0.0046}, "confidence": 0.958, "action": {"act_probability": 1.0}, "usage": {"input_tokens": 37, ...} ``` ### Kotlin — Android GPU with explicit FP32 This is `LayaMinimalUsage.kt` from the sample, compiled into the app (one comment trimmed). It gathers the float16 rows, writes the float32 buffers, runs both graphs and decodes. `LayaEngine` loads the tokenizer, the calibration file and the memory-mapped token table. Build and install steps are in [android/README.md](android/README.md). ```kotlin import android.content.Context import androidx.annotation.WorkerThread object LayaMinimalUsage { /** Call off the main thread after installing the files listed in README.md. */ @WorkerThread fun classify(context: Context): Map { LayaEngine(context, LayaEngine.Storage.WFP16).use { engine -> val question = linkedMapOf( "type" to "choice", "instructions" to "What does the customer need?", "criteria" to linkedMapOf("refund" to "money returned", "help" to "technical help"), ) val backend = LayaEngine.Backend.GPU // CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32). engine.initialize(backend) val row = engine.prepare("同じ支払いが二重に請求されました。返金をお願いします。", question) val raw = engine.runRaw(row, backend) check(raw.finite) return LayaDecoder.decode(raw.markerLogits, raw.actLogits, row.question, engine.calibration) } } } ``` For the NPU, use `LayaEngine.Backend.NPU`: `LayaEngine` then creates the environment with `DispatchLibraryDir` and `CompilerPluginLibraryDir` set to the app's native library directory, requests `QualcommOptions(htpPerformanceMode = BURST)` for the main graph, and keeps the act head on the CPU. ## Host contract `HOST_CONTRACT.md` is the complete specification. In short: 1. Build the sequence exactly as the upstream `build_sequence` does: ` " question: " option0 option1 … text `. Each option keeps at most 48 tokens. The text is truncated on the right so the whole sequence fits in N. Remember the position of every ``. 2. Pad the ids on the right with id 0 to N. `inputs_embeds [1,N,768]` holds the token-table row of every position, padding included. `attention_mask [1,N]` is 1 for real tokens. `qtype_onehot [1,3]` selects choice, score or yes/no. 3. The main graph returns `token_logits [1,N]` and `pooled_cls [1,768]`. Take the logits at the marker positions, divide by the temperature for that question type and option count, and apply softmax. That is the answer distribution. 4. The act head takes `pooled_cls` and four features of the uncalibrated distribution (top probability, top-two margin, normalized entropy, option count / 255) and returns `act_logits [1,2]`. The sample app ships three upstream question presets unchanged: email triage, support intent and moderation. ## Measured quality and performance Reference: the official `laya` 0.3.4 `predict` on fp32 CPU at checkpoint revision `1c5edc17`, with `max_len` set to the window. The inputs are 44 invented states (21 English, 21 Japanese, 2 mixed) with the upstream presets, the upstream quickstart schema and custom schemas of up to 20 options: 201 question rows, of which 81 are choice or score questions and 120 are yes/no. Probabilities are compared at temperature 1, including the yes/no probability and the act probability. | Device, accelerator | Graph | Same argmax | Max probability difference | Placement | Median ms per question | |---|---|---:|---:|---:|---:| | S26 GPU FP32 | S256 wfp16 | 81/81 | 0.0014 | 1809/1809, one partition | 51.9 | | S26 NPU (Hexagon v81, JIT) | S256 wfp16 | 81/81 | 0.0069 | whole graph, one NPU node | 36.5 | | S26 GPU FP32, previous revision | S256 fp32 | 81/81 | 0.0001 | 1680/1680, one partition | 50.8 | | S26 CPU, previous revision | S256 wfp16 | 81/81 | 0.0014 | | 163.0 | | Desktop CPU | S256 and S512, wfp16 | 81/81 | 0.0014 | | | | Desktop CPU | S256 and S512, fp32 | 81/81 | 0.0001 | | | On the phone the Kotlin tokenizer and prompt builder produced the same token ids and marker positions as the Python reference on all 201 rows. The previous-revision rows were measured on 2026-09-21 with the graphs before the fp16 rewrites, which give the same fp32 results. Desktop CPU is ai-edge-litert 2.1.6 on macOS arm64 with 4 threads. The per-question time covers both graphs, from writing the inputs to the end of output readback: the median of 200 rows after one cold call, debug build, screen on, no other workload, battery at 35–37 °C. The host embedding lookup adds a median of 14.6 ms. In the non-debuggable build of the sample, a five-question email-triage run took 301 ms and 303 ms in two fresh launches, and the app reached Ready 2.3 s after `onCreate`. That start-up is 1.26 s of GPU compilation, 0.76 s of tokenizer loading and a 0.27 s warm-up pass. With the NPU selected, the debug build answered the same run in 257 ms and 268 ms and reached Ready 2.0 s after `onCreate`, of which 0.23 s loaded the compiled NPU graph from the cache. These are samples from one device, not a benchmark across devices or thermal states. ## Calibration The upstream multilingual checkpoint ships with every temperature at 1. `laya_ml_calibration.json` holds temperatures fitted here, one per question type and option-count bucket, on 4,415 labeled English and Japanese examples from public datasets, at the 256-token window. A temperature never changes which option wins. | Bucket | Validation rows | Temperature | ECE at T=1 → shipped | |---|---:|---:|---| | choice, 2 options | 140 | 1.40 | 0.089 → 0.117 | | choice, 3–5 | 150 | 1.36 | 0.154 → 0.087 | | choice, 6–10 | 125 | 1.00 | 0.162 → 0.162 | | choice, 11–20 | 108 | 2.45 | 0.383 → 0.120 | | score, 3–5 levels | 161 | 3.68 | 0.257 → 0.087 | | yes/no | 200 | 3.63 | 0.310 → 0.200 | Fits use the same number of English and Japanese rows. A bucket keeps T=1 when a fitted value makes either language worse by more than 0.05 ECE. That is the case for 6–10 options. For 2 options the fit helps Japanese and costs English 0.03, so the pooled figure rises. This is a generic starting point: refit on your own data before you rely on the probabilities. Sources, licenses and label mappings are in [licenses/CALIBRATION_DATASETS.md](licenses/CALIBRATION_DATASETS.md). ## Limits - The numbers above measure agreement with the upstream model, not task accuracy. The conversion reproduces the upstream answers, including the wrong ones. On five of the Japanese rows the upstream answer differs from the intended reading; one example is a veiled threat that scores 0.0 for `threat`. They are listed in `HOST_CONTRACT.md`. - The act probability was 1.0 on all 201 rows, in the upstream model and in these graphs. Do not treat it as an escalation signal without your own evaluation. - The S512 graphs were checked on desktop CPU only. The calibration was fitted at 256 tokens. - One device was tested: one Snapdragon 8 Elite Gen 5 phone for the GPU and the NPU. The token table is memory-mapped and the app needs about 680 MB of files in its private storage. The NPU libraries add about 40 MB to the APK. ## Provenance, conversion and license - Source: `convaiinnovations/laya`, revision `1c5edc17a7acd8701df6fc341c0d179f1c62c982`, subfolder `multilingual/` (encoder [jhu-clsp/mmBERT-base](https://huggingface.co/jhu-clsp/mmBERT-base), 22 layers, hidden 768, vocabulary 256,000), loaded with `laya` 0.3.4 and `transformers` 5.17.0. - Conversion: `litert-torch` 0.9.3 (`torch` 2.12.1), fixed shapes, fp32. The math is unchanged: host-side embedding lookup, the question type as a one-hot matmul, padding and sliding-window masks as float constants, rotary tables baked per layer type, attention kept at rank 4, the two head layers written out explicitly, exact GELU, and the scorer applied at every position. The 2026-09-29 revision adds the three fp16 rewrites described above, all exact in fp32. - Weight storage: `ai-edge-quantizer` 0.8.0 float16 FLOAT_CASTING on the FULLY_CONNECTED weights only (`wfp16` files). - Verification: LiteRT CompiledModel Python API on desktop CPU, and the CompiledModel Kotlin API on the Galaxy S26 for CPU, GPU and NPU. Every gate is the same argmax plus an absolute probability difference. Correlation was never used as a gate. License: the Laya checkpoint and code are Apache-2.0, and these converted files are released under the same license. mmBERT-base declares MIT. The host code ports logic from `laya`, `transformers` and `tokenizers` (Apache-2.0). License texts are in `licenses/`, and attribution is in `NOTICE`. The calibration datasets keep their own terms; none of their text is included here.