Open Decision DeBERTa-v3-Large for LiteRT

com-kotobalabs/open-jev-deberta-v3-large is a decision model built on DeBERTa-v3-large. It reads one state and several typed questions (choice, score, noul) and returns a calibrated probability distribution for each question from one forward pass. This repository holds it converted to LiteRT with litert-torch 0.9.3. The verified path on desktop is the LiteRT CompiledModel Python API on CPU and on the Apple Metal GPU with explicit FP32 computation (ai-edge-litert 2.1.6, Apple M4 Max, 2026-10-01). On a phone, it is the Kotlin CompiledModel API on the Galaxy S26 GPU with explicit FP32 computation (LiteRT 2.2.0, 2026-10-01 and 2026-10-02). On desktop CPU, the shipped graph (512-token window, float16 weights, float16 word table) and the author's implementation (CPU FP32) picked the same option on all 4,327 questions of 1,809 requests, and no probability differed by more than 0.0016 at temperature 1.05.

On the Galaxy S26 GPU, the same graph picked the same option on all 1,302 questions of 500 requests, with a maximum probability difference of 0.00083 and a median of 697.5 ms per request. The 256-token graph also picked the same option on all 1,074 questions of 419 requests, at a median of 407.6 ms. The Kotlin host in android/sample/ ran the fixture gate on the S26: its ids and spans equal the author's Collator on all 160 requests, and all 420 answers match the author's implementation.

An invented support ticket on the left and the three answers the S512 wfp16 graph returned for it on the right

The ticket is invented. The bars are the probabilities that deberta_v3_large_decision_s512_wfp16.tflite with the float16 table returned for it on desktop CPU.

Other formats: onnx-community/open-jev-deberta-v3-large-ONNX, an ONNX port for Transformers.js with fp32, fp16, q4 and q4f16 weights.

Files

File Bytes Role
deberta_v3_large_decision_s512_wfp16.tflite 813,443,728 Graph for a 512-token window, with float16 weights. The examples below use it. Tested on desktop CPU (1,809 requests), on the Metal GPU with explicit FP32 (100 requests) and on the Galaxy S26 GPU with explicit FP32 (500 requests).
deberta_v3_large_decision_s256_wfp16.tflite 712,780,432 Graph for a 256-token window, with float16 weights. Tested on desktop CPU (the 1,281 requests that fit 256 tokens), on the Metal GPU with explicit FP32 (100 requests) and on the Galaxy S26 GPU with explicit FP32 (419 requests).
deberta_v3_large_decision_s256_npu_wfp16.tflite 712,671,952 Graph for a 256-token window, changed for float16 hardware, with float16 weights: the file for the Qualcomm NPU (see NPU). Tested on desktop CPU and on the Galaxy S26 NPU and GPU (the 143 gate requests that fit 256 tokens).
word_embeddings_fp16.bin 262,348,800 Word table [128100, 1024], little-endian float16, for the host lookup.
tokenizer.json 8,657,170 The source repository's tokenizer.json, unchanged.
decision_litert.py Python host: encodes the request, looks up the table rows, builds the routing inputs, runs the graph and reads out the answers.
conversion/ Conversion, fixture, check and figure scripts, with requirements-lock.txt.
android/sample/ Android sample app (Compose) with the pure-Kotlin host that ran on the Galaxy S26: tokenizer, sequence builder, float16 table lookup, LiteRT call and read-out, plus JVM parity tests and a debug fixture gate.
android/CardSnippet.kt The Kotlin block below, in the package it was compiled in. The full sample is in android/sample/.
REPRODUCE.md How to reproduce the conversion and the checks.
LICENSE, NOTICE Apache License 2.0 text and attribution.
assets/hero.png The figure above.
SHA256SUMS SHA-256 checksums of the files.

In the wfp16 graphs only the weights are float16. Each of the 146 FULLY_CONNECTED weights feeds a DEQUANTIZE operator, and activations stay float32. The 48 relative-position constants of the batch matmuls also stay float32. The npu file has 170 FULLY_CONNECTED weights and 1,906 operators (see NPU), with the inputs and outputs of the S256 file.

The measurements below also use three float32 files that are not in this repository: the S256 and S512 fp32 graphs (1,323,008,720 and 1,423,672,016 bytes) and the float32 table (524,697,600 bytes, a bit-exact copy of the checkpoint's word embeddings).

Every graph holds the 24-layer encoder, the span means and the scoring head, and returns one logit per option slot. The host does the rest. It tokenizes, looks up each token's table row, builds the two routing inputs, applies softmax at temperature 1.05 and reads out the answers.

Minimal usage

Python: desktop CPU

Needs numpy, tokenizers and ai-edge-litert (tested with 2.1.6). Torch, transformers and the author's typed_decisions package are not needed. Run it from the directory that holds the downloaded files and decision_litert.py.

from decision_litert import DecisionModel

model = DecisionModel(
    "deberta_v3_large_decision_s512_wfp16.tflite",
    "word_embeddings_fp16.bin",
    "tokenizer.json",
)
state = (
    "Order 8841 was supposed to arrive on Monday and the tracking page still "
    "says 'label created'. I need it before Friday for a gift."
)
questions = [
    {
        "type": "choice",
        "instructions": "Which team should handle this message?",
        "options": ["billing", "technical support", "shipping", "account access", "sales"],
    },
    {
        "type": "score",
        "instructions": "How frustrated is the customer?",
        "options": ["calm", "mildly annoyed", "frustrated", "angry"],
    },
    {"type": "noul", "instructions": "The customer asks for a refund."},
]
team, frustration, refund = model.decide(state, questions)
print(team["choice"], round(team["confidence"], 3))
print(round(frustration["score"], 3), round(frustration["confidence"], 3))
print(round(refund["noul"], 3))
model.close()

On a Mac CPU it prints the following. The code rounds each value to 3 decimals.

shipping 0.785
1.861 0.642
0.124

shipping wins with probability 0.785. The score answer is the expected zero-based level, so 1.861 falls between "mildly annoyed" and "frustrated". Its confidence, 0.642, is the probability of "frustrated". The refund question gives p(yes) = 0.124. On the same request (80 tokens), the author's decide() on CPU FP32 gives the same three answers, and no probability differs by more than 6.6e-5.

decide() returns one dict per question, and the choice and score answers also carry probabilities, one per option. To run on the desktop GPU, pass accelerator="gpu". It sets explicit FP32 (GpuOptions(enforce_f32=True)), the Metal setting measured below.

Kotlin: Android GPU with explicit FP32

import android.util.Half
import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import java.io.File
import java.io.RandomAccessFile
import java.nio.ByteOrder
import java.nio.channels.FileChannel
import kotlin.math.exp

/** One request = token ids plus, per option, the token range of its text and of its question's text. */
class Span(val start: Int, val end: Int)

class DecisionGpu(dir: File, private val seq: Int = 512, private val slots: Int = 128) : AutoCloseable {
  private val options = CompiledModel.Options(Accelerator.GPU).apply {
    gpuOptions = CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32)
  }
  private val model =
    CompiledModel.create(File(dir, "deberta_v3_large_decision_s${seq}_wfp16.tflite").absolutePath, options, null)
  private val inputs = listOf("inputs_embeds", "attention_mask", "q_routing", "o_routing")
    .associateWith { model.createInputBuffer(it, "serving_default") }
  private val outputs = mapOf("logits" to model.createOutputBuffer("logits", "serving_default"))
  private val table = RandomAccessFile(File(dir, "word_embeddings_fp16.bin"), "r").use {
    it.channel.map(FileChannel.MapMode.READ_ONLY, 0, it.length()).order(ByteOrder.LITTLE_ENDIAN).asShortBuffer()
  }
  private val embeds = FloatArray(seq * 1024)

  /** ids from a tokenizer that matches decision_litert.py; one (question span, option span) pair per option, in order. */
  fun logits(ids: IntArray, pairs: List<Pair<Span, Span>>): FloatArray {
    require(ids.size <= seq && pairs.size <= slots)
    for (p in 0 until seq) {
      val row = (if (p < ids.size) ids[p] else 0) * 1024
      for (c in 0 until 1024) embeds[p * 1024 + c] = Half.toFloat(table.get(row + c))
    }
    val qRouting = FloatArray(slots * seq)
    val oRouting = FloatArray(slots * seq)
    pairs.forEachIndexed { j, (q, o) ->
      for (t in q.start until q.end) qRouting[j * seq + t] = 1f / (q.end - q.start)
      for (t in o.start until o.end) oRouting[j * seq + t] = 1f / (o.end - o.start)
    }
    inputs.getValue("inputs_embeds").writeFloat(embeds)
    inputs.getValue("attention_mask").writeFloat(FloatArray(seq) { if (it < ids.size) 1f else 0f })
    inputs.getValue("q_routing").writeFloat(qRouting)
    inputs.getValue("o_routing").writeFloat(oRouting)
    model.run(inputs, outputs, "serving_default")
    return outputs.getValue("logits").readFloat().copyOf(pairs.size)
  }

  /** Probabilities of one question's options at the shipped temperature 1.05. */
  fun probabilities(logits: FloatArray, from: Int, count: Int): List<Double> {
    val z = (from until from + count).map { logits[it] / 1.05 }
    val top = z.max()
    val e = z.map { exp(it - top) }
    val sum = e.sum()
    return e.map { it / sum }
  }

  override fun close() {
    (inputs.values + outputs.values).forEach { it.close() }
    model.close()
  }
}

Create DecisionGpu once and call logits from one worker thread. logits takes the ids of one request and one (question span, option span) pair per option, in request order, and returns one logit per option. The ids and spans must come from a tokenizer that matches decision_litert.py. The block compiles with LiteRT 2.2.0 (AGP 8.9.1, Kotlin 2.2.21). It has not itself run on a device; the sample host below has.

The full Kotlin host that ran on the Galaxy S26 is in android/sample/, with its own README for download, build and install. Its DecisionModel.kt makes the same LiteRT calls as the block above. Its DecisionTokenizer.kt turns text into ids with tokenizer.json, and its DecisionInputs.kt builds the sequence and the spans.

Host contract

decision_litert.py implements this contract. It follows the Collator in the author's typed_decisions/encoder.py and the read-out in typed_decisions/schema.py.

  1. Tokenize each text with tokenizer.json and no added special tokens, then assemble the ids in this order: [CLS] [STATE] state [Q] instructions [OPT] option [OPT] option ... [Q] ... [SEP]. The state keeps its first 256 tokens. The marker ids are [STATE] 128001, [Q] 128002 and [OPT] 128003. [CLS] is 1, [SEP] is 2 and [PAD] is 0. A noul question always has the options ["no", "yes"]. Record the token span of each question's text and of each option's text as half-open ranges, markers excluded.
  2. Check the limits. A choice question takes 2 to 255 options, and a score question 2 to 10 ordered levels. The ids must fit the graph's window (256 or 512 tokens), and all options of the request must fit the 128 option slots. A request that does not fit raises an error instead of being truncated. Only the state is cut, at 256 tokens.
  3. Pad the ids on the right with id 0 to the window S, and fill the four float32 inputs. inputs_embeds [1,S,1024] holds the table row of every position, padding included. attention_mask [1,S] is 1 for real tokens and 0 for padding. In q_routing [1,128,S], row j is 1/len over the text tokens of option j's question. In o_routing [1,128,S], row j is 1/len over option j's text tokens. Rows past the request's last option stay zero.
  4. Run the serving_default signature. It returns logits [1,1,1,128], one logit per option slot in request order. Ignore the slots past the last option. Split the logits by question and apply softmax at temperature 1.05 within each question's options.
  5. Read the answer. choice is the option with the highest probability, and confidence is that probability. score is Σ i·pᵢ, the expected zero-based level, which may fall between levels. It also carries confidence. noul is p(yes).

On all 1,809 fixture requests, decision_litert.py produces the same ids and text spans as the author's Collator. Its decide() also gives the same answers as the author's decide() on 209 requests: the README example, the 8 invented tickets and the first 200 requests from data-fam. With the S512 wfp16 graph and the float16 table on desktop CPU, probabilities differ by at most 0.0016 and score values by at most 0.00066. With the S256 wfp16 graph, probabilities differ by at most 0.00081.

The tokenizers library on tokenizer.json gives the same ids as the author's AutoTokenizer on all 2,368 fixture texts. The SentencePiece spm.model gives the same ids on 2,367 of them, so build the ids from tokenizer.json.

Measured quality and performance

The reference is the author's implementation: the checkpoint's own typed_decisions package on CPU FP32 (torch 2.12.1, transformers 4.57.6, eager attention, 8 threads), one request per forward pass. Probabilities are compared at temperature 1.05. "Same argmax" means the same winning option for a question.

The fixtures are 1,809 requests with 4,327 questions. They are the first 1,500 states of the author's public data-fam/test.jsonl, the first 300 states of data-fam/ood-test.jsonl, the author's README example and 8 invented tickets with 4 questions each. Both test files come from the author's kotoba-lang/typed-decisions repository at revision e65215db7e554a31f557936f0bd1d8a0a78b6af5. A boundary question is one where the reference's top probability is below 0.9. There are 1,830 of them among the 4,327 questions, and 1,602 among the 2,796 questions that fit 256 tokens.

The desktop rows ran on 2026-10-01 on one desktop: Apple M4 Max (16 cores, 128 GB), macOS 27.0, ai-edge-litert 2.1.6 through the Python CompiledModel API. The Galaxy S26 rows ran on 2026-10-01 and 2026-10-02 on one Galaxy S26 (SM-S942Q, Android 16) with LiteRT 2.2.0 through the Kotlin CompiledModel API. They used a debug gate app, except the row marked as the sample in android/sample/. Times are informational. Each CPU row states the load on the machine. Each phone row states the thermal status and battery temperature where they were recorded, and the compile time. Phone times are warm medians of writing the inputs, running and reading back, with [min, max] where recorded.

Where Graph, word table Requests Questions (boundary) Same argmax Max probability difference Questions over 0.01 Median ms per request
CPU, 8 threads, 3 jobs in parallel S256 fp32, float32 table 1,281 2,796 (1,602) 2,796 (1,602) 1.1e-5 0 152.8
CPU, 8 threads, 3 jobs in parallel S256 wfp16, float16 table 1,281 2,796 (1,602) 2,796 (1,602) 0.00081 0 305.8
CPU, 8 threads, 2 to 3 jobs in parallel S256 fp32, float16 table 1,281 2,796 (1,602) 2,796 (1,602) 0.00024 0 135.9
CPU, 8 threads, 2 to 3 jobs in parallel S256 wfp16, float32 table 1,281 2,796 (1,602) 2,796 (1,602) 0.00080 0 350.2
CPU, 8 threads, 3 jobs in parallel S512 fp32, float32 table 1,809 4,327 (1,830) 4,327 (1,830) 1.1e-5 0 291.1
CPU, 8 threads, 2 to 3 jobs in parallel S512 wfp16, float16 table (shipped) 1,809 4,327 (1,830) 4,327 (1,830) 0.0016 0 531.8
CPU, 8 threads, 2 jobs in parallel S512 fp32, float16 table 1,809 4,327 (1,830) 4,327 (1,830) 0.00024 0 214.2
CPU, 8 threads, 1 to 2 jobs in parallel S512 wfp16, float32 table 1,809 4,327 (1,830) 4,327 (1,830) 0.0016 0 467.8
CPU, idle machine, 20-request probe S512 fp32, float32 table 20 48 48 1.7e-6 0 203.6
Metal GPU, explicit FP32 S512 fp32, float32 table 100 222 222 3.5e-6 0 68.4
Metal GPU, explicit FP32 S512 wfp16, float16 table (shipped) 100 222 (92) 222 (92) 0.0016 0 68.2
Metal GPU, explicit FP32 S256 wfp16, float16 table 100 200 200 0.00081 0 27.6
Metal GPU, default precision (fp16 kernels) S512 wfp16, float16 table 100 222 0: every output non-finite 55.3
Metal GPU, default precision (fp16 kernels) S256 wfp16, float16 table 100 200 0: every output non-finite
Metal GPU, default precision (fp16 kernels) S512 fp32, float32 table 100 222 0: every output non-finite 55.9
Galaxy S26 GPU, explicit FP32; thermal status 0 to 3, battery 31.0 to 44.9 °C; compile 3.69 s S512 wfp16, float16 table (shipped) 500 1,302 (754) 1,302 (754) 0.00083 0 697.5 [575.5, 1,201.7]
Galaxy S26 GPU, explicit FP32; thermal status 2, battery 42.4 to 43.0 °C; compile 4.97 s S256 wfp16, float16 table (shipped) 419 1,074 (712) 1,074 (712) 0.00074 0 407.6 [375.6, 484.7]
Galaxy S26 GPU, explicit FP32; thermal status 2 on the debug-gate runs before and after; the sample in android/sample/ with its Kotlin tokenizer Shipped S256 (143 requests) and S512 (17 requests), float16 table 160 420 420; ids and spans equal to the author's Collator on 160 of 160 requests 0.00069 0 301 (S256 299, S512 620)
Galaxy S26 GPU, default precision (fp16); thermal status 2, battery 43.0 to 42.3 °C; compile 5.12 s S512 wfp16 with a -1e4 mask (experiment), float16 table 100 274 (189) 216 (142): finite but wrong 0.666 248 1,011.9 [933.5, 1,064.3]
Galaxy S26 NPU (Hexagon, JIT, BURST mode) S512 wfp16, float16 table (shipped) none compile failed: the HTP prepare ran out of host memory 94 s after the compile started
Galaxy S26 NPU (Hexagon, JIT, BURST mode); thermal status 1 to 2, battery 40.8 to 41.9 °C; JIT compile 186.9 s S256 wfp16, float16 table (shipped) 100 276 (192) 95 (54): finite but wrong 0.9986 265 129.7 [124.2, 140.3]
Galaxy S26 NPU (Hexagon, JIT, BURST mode); thermal status 2, battery 40.9 to 42.4 °C; JIT compile 200.8 s S256 wfp16 with a -1e4 mask (experiment), float16 table 100 270 (187) 102 (65); 6 more questions non-finite 0.999 259 128.3
Galaxy S26 NPU (Hexagon, JIT, BURST mode); JIT compile 180.7 s S256 wfp16 with a -1e4 mask and SafeLayerNorm k=2 (experiment), float16 table 100 268 112; 8 more questions non-finite 0.998 258 111.8
Galaxy S26 NPU (Hexagon, JIT, BURST mode); JIT compile 171.8 s S256 wfp16 with a -1e4 mask and SafeLayerNorm k=3 (experiment), float16 table 100 276 (192) 274 (190), all finite 0.0254 7 110.0 [106.9, 117.4]
Galaxy S26 NPU (Hexagon, JIT, BURST mode); JIT compile 182.5 s S256 wfp16 with a -1e4 mask and SafeLayerNorm k=4 (experiment), float16 table 100 276 (192) 273 (189), all finite 0.0279 11 108.7 [102.6, 112.8]
Galaxy S26 CPU; Pixel phones: GPU, NPU, CPU any not measured

S512 and S256 name the window size. The S256 CPU rows cover the 1,281 requests that fit 256 tokens. On the shipped S512 CPU row, every question kind matches: choice 1,802 of 1,802, score 716 of 716 and noul 1,809 of 1,809. The largest change in a score answer, the expected level, is 0.0009. The CPU probe compiled in 0.5 s, and the shipped S512 graph compiled in 3.25 s on Metal with explicit FP32.

The S26 S512 row with explicit FP32 uses 500 of the fixture requests: all 300 from ood-test.jsonl, the README example, the 8 invented tickets and 191 random requests from test.jsonl. The S26 S256 row covers the 419 of them that fit 256 tokens. On both rows the largest change in a score answer is 0.00079. The sample row uses 160 fixture requests: the README example, the 8 invented tickets, 100 from ood-test.jsonl and 51 from test.jsonl.

With explicit FP32 on the S26 GPU, each shipped graph ran all 1,785 operators in one LITERT_CL partition. The initial call took 620.9 ms at S512 and 428.1 ms at S256. The Kotlin float16 table lookup is not in these times: it took a median of 37.5 ms per S512 request and 21.7 ms per S256 request in the debug gate app. On the NPU, the shipped S256 graph and the k=3 graph each ran whole as one DispatchDelegate node. Where a row has non-finite outputs, its question count covers only the requests whose outputs were finite.

At the default Metal precision, every output is non-finite on all three graphs tried. On the S26 NPU, the shipped S256 graph gives finite but wrong answers. The next section explains both.

Accuracy on the author's public test file comes from the author's implementation at temperature 1.05, scored with the author's metrics code. On the first 1,500 states of data-fam/test.jsonl (3,474 questions), accuracy is 0.8584. By domain it is 0.9531 for banking77, 0.7560 for sst5 and 0.8752 for boolq, and score questions reach 0.5726. On the first 300 states of data-fam/ood-test.jsonl (818 questions), accuracy is 0.7017, with 0.4312 on score questions. The metric is argmax accuracy, and on desktop CPU the shipped graph picks the same option on every one of these questions, so it scores the same.

The author reports 0.854 in-domain and 0.690 out of domain. Those are the author's numbers, on a test set of 1,500 states and 3,508 questions (4,012 out-of-domain questions) that is not in the author's repository. The repository's data-fam/test.jsonl is a later generation, with 489 sst5 states against 506 and 513 boolq states against 496. Our numbers are on that public file, not on the author's test set.

Not measured:

  • Pixel phones: GPU, NPU and CPU.
  • The Galaxy S26 CPU.
  • Sustained runs.

The NPU rows above ran the two default graphs and four S256 experiments. The NPU section describes the file this repository ships for the NPU.

fp16, the GPU default precision and the NPU

Use the GPU with explicit FP32 computation. In fp16 computation, two parts of the graph fail.

One is the attention mask constant. The graph uses torch.finfo(float32).min, -3.4e38, the stock transformers DeBERTa value. In fp16 this constant becomes -inf. At every real token (1 - mask) is 0, and 0 × -inf is NaN. At the default Metal precision, which runs fp16 kernels, every output is non-finite, for the fp32 graph and the wfp16 graphs alike. A torch fp16 run of the graph agrees: the earliest non-finite module is layers.0.layer.attention.output.dense, and 0 logits are finite. On the S26 NPU, the shipped S256 graph returns finite outputs, but only 95 of 276 answers match the author's implementation.

The other is LayerNorm. Apart from the mask, the fp32 values are small. On the 413-token / 89-option request, the residual stream peaks at 28.5, the attention scores at 32.4 and the FFN intermediate at 142. But LayerNorm sums squares over 1024 channels, and 1024 × 28.5² is 8.3e5, above the fp16 maximum of 65,504. A -1e4 mask alone keeps fp32 results bit-identical on 6 probe requests at S512. On Metal at default precision the outputs then become finite, but only 185 of 222 answers match. The S26 GPU at default precision matches 216 of 274, and the S26 NPU 102 of 270, with 6 more questions non-finite.

The fix has two parts, and it is exact in fp32: the -1e4 mask, and SafeLayerNorm with a shift of k=3, which computes all 49 LayerNorms on x·2⁻³ with eps·2⁻⁶. The graph stays bit-identical in fp32 on 5 probe requests. With both parts, Metal at default precision matches all 200 answers of 100 S256 requests, with a maximum probability difference of 0.0152. The S26 NPU matches 274 of 276 answers, with a maximum probability difference of 0.0254, at a median of 110.0 ms per 256-token request. The S26 GPU with explicit FP32 takes 407.6 ms for the same window. The two answers that differ are near-ties: the reference's top two probabilities differ by 0.0065 and 0.040. A shift of k=2 still overflows, with 8 questions non-finite. A shift of k=4 is no better, at 273 of 276.

The two default graphs ship unchanged: neither rewrite is applied to them. Run them on the GPU with explicit FP32. The k=3 NPU graph is not included, because two of its answers differ from the author's implementation. The NPU section describes the S256 graph that ships for the NPU instead. The S512 graph cannot be compiled for the NPU on the S26: the HTP prepare ran out of host memory 94 s after the compile started.

Explicit FP32 is the GPU setting verified here: GpuOptions(enforce_f32=True) on Metal and GpuOptions(precision = FP32) on the S26.

NPU

deberta_v3_large_decision_s256_npu_wfp16.tflite runs on the Qualcomm NPU of the Galaxy S26 (Hexagon v81 HTP). On the 143 gate requests that fit 256 tokens, it picked the same option as the author's implementation on all 371 questions. The default S256 file also runs on that NPU, but it picked the same option on 95 of 276 questions (the table in Measured quality and performance).

The NPU computes in float16. Three changes keep the graph within float16's range and precision. Each is exact in float32. In PyTorch, the changed graph's logits equal the default S256 graph's bit for bit on 5 probe requests. On desktop LiteRT CPU, the npu file and the default S256 wfp16 file return bit-identical outputs on all 143 gate requests.

  • Attention mask. Masked positions add −1e4 instead of float32's lowest value, which becomes −inf in float16 (the section above).
  • LayerNorm pre-scale. Each of the 49 LayerNorms computes on x·2⁻ᵏ with eps·4⁻ᵏ. In float32, the inputs of the 24 feed-forward output norms reach |x| = 1,527 before normalization. Their squared deviation from the mean reaches 2.3e6 for one element and 4.9e6 summed over a row, above the float16 maximum of 65,504. For each of these norms, k is the smallest value of at least 1 that keeps the largest squared deviation at or under 4,096 and the largest row sum at or under 16,384 on 145 requests. The values run from 1 to 5. The embedding norm and the 24 attention-output norms use k = 1. Their row sums stay under 2,600. conversion/ln_k_d1.json lists every k.
  • Feed-forward split. The first FULLY_CONNECTED of each feed-forward block runs as two halves joined by CONCATENATION before the GELU. On the k=3 graph of that table and the 143 gate requests, the NPU picked the same option on 368 of the 371 questions without the split and on all 371 with it.
Galaxy S26, LiteRT 2.2.0 File Same argmax (371 questions) Max probability difference Median ms per request
NPU (JIT, BURST) npu 371/371 0.0093 119.4
GPU, explicit FP32 npu 371/371 0.00069 184.3
GPU, FP16 with FP32 accumulation npu 371/371 0.0058 145.9

Each row is one process on 2026-10-05, started at thermal status 0, in the device gate of hfmodels-android (a Kotlin host). The time includes the float16 table lookup: it is the median over the 143 requests of lookup, input write, run and readback. The questions of a request share one forward, as in the author's decide(). The first load compiles the graph for the NPU on the phone: 268.1 s. A second process read LiteRT's compiled cache in 1.22 s and picked the same options. With explicit FP32 on the GPU, the npu file gave the same maximum difference as the default S256 file on these requests, 0.00069. The phone went from thermal status 0 to 2 during each GPU run.

The NPU needs Qualcomm's runtime inside the app: the dispatch library and JIT compiler plugin of the LiteRT 2.2.0 release, and the QAIRT 2.47 HTP libraries for Hexagon v81. In Kotlin, create the Environment with the app's native library directory as DispatchLibraryDir and CompilerPluginLibraryDir, and pass Accelerator.NPU with QualcommOptions(htpPerformanceMode = BURST). Every NPU run above used BURST. The Android sample in android/sample/ runs the GPU path only.

Limits

  • The agreement numbers measure how closely the conversion follows the author's implementation, not task accuracy. The conversion reproduces the author's answers, including the wrong ones.
  • English only, per the author. Training covers three public domains: banking77 support messages, SST-5 review sentences and BoolQ passages. The author says anything else is out of distribution and should be measured before use.
  • The window is 512 tokens, and the author's contract cuts the state to 256 tokens. Of the 1,809 fixture requests, 1,281 fit the S256 graph and all fit the S512 graph. The longest has 485 tokens.
  • A request has 128 option slots, shared by all its questions. The author's contract allows up to 255 options in one choice question, and such a request does not fit this graph. The fixtures need at most 89.
  • A request that does not fit the window or the option slots raises an error. Only the state is cut.
  • In the author's words, the model "reads the question only partly". The author reports a gap between in-domain and out-of-domain questions, and "unseen ordered scales are the weakest". On the public test file the gap is 0.8584 against 0.7017, and score questions out of domain reach 0.4312.
  • Confidence comes from a temperature fitted on the author's validation split. The author reports an in-domain ECE of 0.022 and advises re-calibrating on your own data for out-of-domain questions.
  • The model does not generate text. It chooses among the options you give it.
  • One desktop and one phone, the Galaxy S26, were tested. The Kotlin tokenizer is in android/sample/.
  • NPU validation covers the Galaxy S26 (Hexagon v81) only. Only the 256-token window has an NPU file: the S512 graph did not compile for the NPU, and the default S256 file picks wrong options on it. The first NPU load on a phone compiles the graph for 268.1 s; later loads read LiteRT's cache.
  • The phone GPU warms under load. Over the 500-request S512 run, the thermal status rose from 0 to 3, and later requests took longer than early ones. Sustained throughput is lower than the early requests show.

Provenance, conversion and license

  • Source: com-kotobalabs/open-jev-deberta-v3-large at revision 188ee67a5c93122b916e5acd5bdb0cb3623e380a. model.safetensors is 1,736,094,384 bytes, all F32, with 434,012,160 parameters, SHA-256 3f1d5bc3b6d3dc412ea2c499446fcbd212242d79baf16bc0eda155b4ebbac806. head.safetensors is 12,591,412 bytes, SHA-256 f101be67c5808a810abbcede9ebe702f7e4af4aa6f7298bb243842ce9e3a295a.
  • Model: a DebertaV2Model encoder with 24 layers, hidden size 1024, 16 heads, FFN size 4096, relative attention with 256 position buckets, no absolute position embeddings and a vocabulary of 128,100. For each option, the head applies a 3072-to-1024 linear layer, GELU and a 1024-to-1 linear layer to [mean of the question's text tokens; mean of the option's text tokens; their product]. Softmax runs within each question's options at temperature 1.05, which the author fitted on the validation split.
  • Training data, per the author: public gold labels only, from banking77 (CC-BY-4.0), SST-5 (the author names no license) and BoolQ (CC-BY-SA-3.0). The base model is microsoft/deberta-v3-large, which the author lists as MIT.
  • The author describes the model as an "independent reproduction of the shape" of a hosted decision API. It is "not affiliated" with that API's maker and uses none of its data or code.
  • Conversion: litert-torch 0.9.3 with torch 2.12.1 and transformers 4.57.6, at fixed shapes. The encoder layers are a DeBERTa-v2 layer rewritten for fixed shapes, in conversion/shaped_layer.py (SHA-256 prefix 1adada39), the same file used for the GLiNER2.5 conversion. It pre-expands the log-bucket relative positions as constants and uses rank-4 batch matmuls, a float mask and native GELU.
  • The author's span pooling becomes the two routing inputs, q_routing and o_routing. The question and option means are then matrix products inside the graph.
  • Each graph has 1,639 operators, or 1,785 in the wfp16 form; the npu graph has 1,736, or 1,906 in its wfp16 form. None is GATHER, GATHER_ND, CAST, SELECT_V2, BROADCAST_TO or MAXIMUM, and no BATCH_MATMUL has a constant left operand. No tensor is int64, and none is above rank 4.
  • Float16 weights: ai-edge-quantizer 0.8.0, FLOAT_CASTING, weight-only.
  • fp16 experiments: conversion/graph.py carries the two optional fp16 rewrites used in the experiments, the finite mask constant and SafeLayerNorm. conversion/export.py turns them on with --mask-neg and --ln-shift. The two default graphs use neither.
  • NPU file: conversion/npu_graph.py holds its three changes. conversion/export_npu.py checks them in PyTorch fp32, exports the graph with the k of conversion/ln_k_d1.json and casts the weights to float16 with the same quantizer recipe. Step 4 of REPRODUCE.md rebuilds the published file byte for byte.
  • Rewrite check: in PyTorch fp32, the rewritten graph is bit-identical to the stock DebertaV2Model path. The maximum difference is 0.0 on 5 probe requests at S256 and 6 at S512, including the 485-token / 89-option request. The rewritten graph's logits are within 9.5e-7 of the author's implementation.
  • Word table: the checkpoint's float32 word embeddings rounded to the nearest float16, with no inf or NaN. SHA-256 d1b86bce8ffa0aae67e7d99db5824e67b5241337d0b5132bf25f22bf9dd72b45. tokenizer.json has SHA-256 cd119378b0160677b7a1e561ba29ada83918c1d420b326516a585382e83d9d39.
  • Verification: the LiteRT CompiledModel Python API on desktop CPU and Metal, and the Kotlin CompiledModel API on the Galaxy S26. Every check is the same argmax plus the absolute probability difference against the author's implementation.

License: the author's model artifacts are licensed under Apache 2.0. These converted files, the conversion scripts, the Python host and the Kotlin block are released under the same license. DeBERTa-v3-large is MIT. The source model is published by Mithril (formerly Kotoba Cloud), operated by Kotoba Labs Inc. The license text is in LICENSE, and attribution is in NOTICE.

Downloads last month
21
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for litert-community/Open-Decision-DeBERTa-v3-Large-LiteRT

Finetuned
(3)
this model