zlm-v1-signal-extract

A 77M-parameter FLAN-T5 model that reads a short text and returns its keywords and the user intent behind it β€” in one generative pass, in ~136 ms on CPU.

Given an ad headline, search query, or product blurb, the model generates two signals that normally require two separate systems (a keyword extractor and an intent classifier):

  • keywords β€” the terms that carry the text's meaning, and
  • user intent β€” a free-text intent phrase plus one of three standard search-intent classes: informational, transactional, or navigational.

On an independent 1,000-example benchmark it beats a frontier-LLM baseline (GPT-5-nano) on both jobs β€” +0.10 semantic keyword F1 and +0.32 intent-category accuracy β€” while running ~9Γ— faster. It is built for high-volume, latency-sensitive enrichment (contextual ad targeting, query understanding, analytics) where an LLM call per request is too slow and too expensive.

The repository ships the model in two layouts built from the same weights: the server layout (split decoders, for Optimum / onnxruntime) and the on-device layout (merged decoder, for transformers.js in browsers, Node.js and workers), the same bundle ZeroGPU runs on edge devices in production.

How it works

The model is a fine-tuned google/flan-t5-small (encoder-decoder, 77M parameters). Both signals come out as a single short sequence:

Input:   extract keywords and user intent: <your text>
Output:  keyword 1, keyword 2, ... [SEP] <intent name> [SEP] <intent category>

Example (actual greedy output of the published files):

Input:   extract keywords and user intent: Sign up for our newsletter and get 20% off your first order of organic skincare products.
Output:  newsletter, 20% off, organic skincare products, free shipping, promo, e-mail [SEP] sign up for newsletter [SEP] transactional

It reads up to 256 input tokens and generates up to 48. Input is expected in English β€” in multilingual deployments, translate to English first.

Files

Mixed precision throughout: the encoder is fp32 and the decoders are int8 (dynamic quantization).

Server layout (standard Optimum split-ONNX):

File Precision Role
encoder_model.onnx fp32 Runs once per input
decoder_model_quantized.onnx int8 (dynamic) First generation step
decoder_with_past_model_quantized.onnx int8 (dynamic) Subsequent steps (KV-cache)

On-device layout (transformers.js convention, loaded with dtype: 'q8'):

File Precision Role
onnx/encoder_model_quantized.onnx fp32 (see note) Runs once per input
onnx/decoder_model_merged_quantized.onnx int8 (dynamic) All generation steps (use_cache_branch switch)
postprocess.json β€” Serving defaults: prompt template, [SEP] token, token limits, model version

Both layouts share config.json, generation_config.json and the tokenizer files (tokenizer.json, spiece.model, tokenizer_config.json, special_tokens_map.json, added_tokens.json).

Notes on the on-device layout:

  • onnx/encoder_model_quantized.onnx is fp32 despite its name β€” it is byte-identical to encoder_model.onnx. transformers.js uses a single quantized flag for both seq2seq graphs, so the fp32 encoder carries the quantized file name.
  • In production this layout is served to devices with at least 8 GB of RAM, with a 256-token input cap.
  • The merged decoder was built from the two int8 server decoders with Optimum's merge_decoders; its greedy output matches the split decoders token-for-token on the same onnxruntime build. It is saved at ONNX IR version 8 (opset 18) so it also opens in older onnxruntime-web builds.

Why mixed precision: quantizing the encoder collapses output quality (measured keyword F1: fp32 encoder 0.451 β†’ full-int8 0.044), because every generated token conditions on the encoder's hidden states. The encoder runs once, so fp32 there is cheap; the decoders run per output token, so int8 there is where the speed comes from.

Usage

Python (Optimum + onnxruntime)

With optimum[onnxruntime]:

from optimum.onnxruntime import ORTModelForSeq2SeqLM
from transformers import AutoTokenizer

repo = "ZeroGPU/zlm-v1-signal-extract"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = ORTModelForSeq2SeqLM.from_pretrained(
    repo,
    encoder_file_name="encoder_model.onnx",
    decoder_file_name="decoder_model_quantized.onnx",
    decoder_with_past_file_name="decoder_with_past_model_quantized.onnx",
)

def extract(text: str):
    inputs = tokenizer(f"extract keywords and user intent: {text}", return_tensors="pt", truncation=True, max_length=256)
    out = model.generate(**inputs, max_new_tokens=48)
    decoded = tokenizer.decode(out[0], skip_special_tokens=True)
    parts = [p.strip() for p in decoded.split("[SEP]")]
    keywords = [k.strip() for k in (parts[0] if parts else "").split(",") if k.strip()]
    return {
        "keywords": list(dict.fromkeys(keywords))[:8],  # dedupe + cap, as in production
        "intent_name": parts[1] if len(parts) > 1 else "",
        "intent_category": parts[2] if len(parts) > 2 else "",
    }

print(extract("Best price on wireless noise-cancelling headphones"))

Pass truncation=True, max_length=256 explicitly as above: the tokenizer files carry a 64-token truncation default, and the model is served with a 256-token input window.

JavaScript (transformers.js v3 β€” browser, Node.js, workers)

import { AutoModelForSeq2SeqLM, AutoTokenizer, Tensor } from '@huggingface/transformers';

const repo = 'ZeroGPU/zlm-v1-signal-extract';
const tokenizer = await AutoTokenizer.from_pretrained(repo);
// dtype 'q8' selects the on-device layout: onnx/encoder_model_quantized.onnx + onnx/decoder_model_merged_quantized.onnx
const model = await AutoModelForSeq2SeqLM.from_pretrained(repo, { dtype: 'q8' });

async function extract(text) {
  // 256-token input window; keep the closing </s> (transformers.js' own `truncation` option drops it)
  let ids = Array.from(tokenizer(`extract keywords and user intent: ${text}`).input_ids.data, Number);
  if (ids.length > 256) ids = [...ids.slice(0, 255), ids.at(-1)];
  const tensor = (data) => new Tensor('int64', BigInt64Array.from(data, BigInt), [1, data.length]);
  const out = await model.generate({ input_ids: tensor(ids), attention_mask: tensor(ids.map(() => 1)), max_new_tokens: 48 });
  const [decoded] = tokenizer.batch_decode(out, { skip_special_tokens: true });
  const [keywords = '', intentName = '', intentCategory = ''] = decoded.split('[SEP]').map((s) => s.trim());
  return {
    keywords: [...new Set(keywords.split(',').map((k) => k.trim()).filter(Boolean))].slice(0, 8),
    intentName,
    intentCategory,
  };
}

console.log(await extract('Best price on wireless noise-cancelling headphones'));

Notes:

  • [SEP] is registered as a non-special token, so it survives skip_special_tokens=True decoding β€” that is what makes the three-field output parseable.
  • The int8 decoders are sensitive to the onnxruntime build: on a near-tie between two tokens, different builds (for example onnxruntime 1.30 in Python and onnxruntime-node 1.21 / 1.24) can pick different tokens, and greedy decoding then continues from there. In our checks this changed 4 of 10 sample outputs between Python and Node, while the split and merged decoders agreed exactly within each runtime. Expect small output differences across runtimes, not across the two layouts.
  • generative_enrichment_metadata.json (server layout) and postprocess.json (on-device layout) carry the same serving defaults.

Evaluation

Two complementary benchmarks: a balanced independent held-out set, and a harder real-world multilingual stress test.

Independent 1,000-example benchmark (vs GPT-5-nano)

Metric zlm-v1-signal-extract GPT-5-nano
Keyword F1 (semantic match) 0.566 0.467
Keyword F1 (exact match) 0.381 0.235
Intent category accuracy 0.915 0.594
Intent name ROUGE-1 0.327 0.358
Latency p50 (CPU) 136 ms 1,218 ms

Per-class intent accuracy: informational 0.895, transactional 0.887, navigational 0.975 β€” the model discriminates all three classes rather than defaulting to the majority one.

5,000 real-world multilingual prompts (vs GPT-5.4-nano, GPT-5.5 gold)

Keyword F1 zlm-v1-signal-extract GPT-5.4-nano
All 5,000 prompts 0.465 0.415
English 0.479 0.609
Non-English (translated) 0.456 0.286

The model's keyword quality is essentially language-stable (0.479 vs 0.456), while the LLM baseline loses more than half its accuracy outside English. On this production-skewed set (~69% informational) the model's intent calibration is its honest weak spot: category accuracy 0.527 vs the LLM's 0.845 (always answering the majority class would score 0.691) and macro-F1 0.381 vs 0.628, driven by over-predicting the rare navigational class. Trained class-balanced, it is excellent on balanced data (0.915 above) and miscalibrated under heavy skew β€” recalibrating class priors to the deployment distribution is the known next improvement.

Both benchmarks were run on the server layout. The on-device layout uses the same weights and the same int8 decoders (merged), so the numbers apply to it as well, within the runtime caveat in Usage.

Training

  • Fine-tuned from google/flan-t5-small on a teacher-labelled corpus of short real-world texts; keywords and intent names were relabelled by GPT-5-nano (the same model used as the baseline in the first benchmark below), with the rare navigational class oversampled (from under 1% to roughly a third of examples) so all three intent classes are learned.
  • Trained in fp32: FLAN-T5 was pretrained in bf16 and overflows to NaN loss under fp16 β€” a practical caveat for anyone reproducing the fine-tune on pre-Ampere GPUs.
  • Exported to split ONNX (encoder / decoder / decoder-with-past) with dynamic int8 quantization on the decoders only, for the reasons in Files; the on-device merged decoder is built from those two int8 decoders.

The training and evaluation datasets are not published.

Limitations

  • English input only β€” quality degrades on untranslated non-English text.
  • Intent calibration under skew β€” see Evaluation; apply your own class priors or thresholds if your traffic is heavily informational.
  • The intent name is free-form generated text; treat it as a descriptive hint, not a controlled vocabulary.
  • Keywords can repeat within one output (for example overtime, overtime); dedupe them as in the Python example. They can also include related terms that do not appear literally in the input (free shipping in the example above).
  • Benchmark gold labels are LLM judgements (GPT-5.5 / independent labelling), not human annotation β€” read absolute numbers as relative comparisons between systems.
  • Output is capped at 48 tokens: roughly 6–10 keywords plus the intent fields.

Changelog

  • 2026-09-29 β€” Added the on-device layout that ZeroGPU serves to edge devices: onnx/encoder_model_quantized.onnx (fp32, byte-identical to encoder_model.onnx), onnx/decoder_model_merged_quantized.onnx and postprocess.json. tokenizer.json gained add_prefix_space on its Metaspace step for transformers.js 2.x (token ids unchanged; verified on 3,000+ texts). generation_config.json now sets max_new_tokens: 48 and writes eos_token_id as a scalar. Weights are unchanged (same flan-t5-small-v2-teacher-relabel-mixed checkpoint), and every file of the server layout is untouched. Card corrections: named the relabelling teacher (GPT-5-nano) and added the 5k-set intent accuracy.
  • 2026-08-07 β€” Initial release (server layout).

License and attribution

  • Weights: Apache-2.0 (same as the base model).
  • Built on FLAN-T5 (Google).

Citation

@misc{zerogpu2026signalextract,
  title  = {zlm-v1-signal-extract: single-pass keyword and user-intent extraction with a fine-tuned FLAN-T5-small},
  author = {ZeroGPU},
  year   = {2026},
  url    = {https://huggingface.co/ZeroGPU/zlm-v1-signal-extract}
}
Downloads last month
1,045
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for ZeroGPU/zlm-v1-signal-extract

Quantized
(25)
this model

Evaluation results

  • Keyword F1 (semantic) on Independent 1,000-example held-out benchmark (LLM-labelled)
    self-reported
    0.566
  • Keyword F1 (exact) on Independent 1,000-example held-out benchmark (LLM-labelled)
    self-reported
    0.381
  • Intent category accuracy on Independent 1,000-example held-out benchmark (LLM-labelled)
    self-reported
    0.915
  • Keyword F1 (overall) on 5,000 real-world prompts, 31 languages (GPT-5.5 gold)
    self-reported
    0.465