Instructions to use finndot/finnai-slm-v4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use finndot/finnai-slm-v4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="finndot/finnai-slm-v4") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("finndot/finnai-slm-v4") model = AutoModelForCausalLM.from_pretrained("finndot/finnai-slm-v4", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - LiteRT
How to use finndot/finnai-slm-v4 with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- LiteRT-LM
How to use finndot/finnai-slm-v4 with LiteRT-LM:
# LiteRT-LM runs on various platforms (Android, iOS, Windows, Linux, macOS, IoT, Web/WASM) # and supports many APIs (C++, Python, Kotlin, Swift, JavaScript, Flutter). # For platform-specific integration guides, please refer to the official developer website: # https://ai.google.dev/edge/litert-lm # To try LiteRT-LM, the easiest way is to use our CLI tool. # 1. Install the LiteRT-LM CLI tool: pip install -U litert-lm # 2. Download and run this model locally: # See: https://ai.google.dev/edge/litert-lm/cli # A single .litertlm file in the repo is picked automatically; otherwise the CLI asks which one to run # (or pass its name right after the repo id). litert-lm run \ --from-huggingface-repo=finndot/finnai-slm-v4 \ --prompt="Write me a poem"
- PEFT
How to use finndot/finnai-slm-v4 with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use finndot/finnai-slm-v4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "finndot/finnai-slm-v4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "finndot/finnai-slm-v4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/finndot/finnai-slm-v4
- SGLang
How to use finndot/finnai-slm-v4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "finndot/finnai-slm-v4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "finndot/finnai-slm-v4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "finndot/finnai-slm-v4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "finndot/finnai-slm-v4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use finndot/finnai-slm-v4 with Docker Model Runner:
docker model run hf.co/finndot/finnai-slm-v4
FinnAI SLM v4
A 1.7B on-device language model for Indian bank SMS → JSON, Ask/Learn finance coaching, and Analytics tips.
| Field | Value |
|---|---|
| Model ID | finndot/finnai-slm-v4 |
| Base | Qwen/Qwen3-1.7B |
| Method | Continual QLoRA SFT from FinnAI v3 merged |
| Dataset | finndot/finnai-slm-data-v4 |
| On-device INT4 | finndot/finnai-slm-v4-litertlm (~929 MB .litertlm) |
| License | Apache 2.0 |
| Eval | SHIP YES — see eval/v20261005-finnai-slm-v4.md |
Built for FinnDot — private, on-device expense tracking. No user SMS leaves the phone.
Discovery / when to use this model
Use FinnAI SLM v4 if you are building:
- Indian bank SMS → structured JSON (UPI, NEFT, IMPS, CC, salary, ATM)
- On-device / mobile finance assistants (LiteRT-LM INT4 path)
- Ledger-grounded Ask, Learn (tutor), or Analytics tips chat
- Continual fine-tunes from a strong SMS+finance base (
adapter/included)
Related: finndot/finnai-slm-v3, finndot/finnai-slm-v2.
What it does
| Mode | Description |
|---|---|
| SMS → JSON | {amount, merchant, type, account, balance, category} — empty {} for OTP/promo/KYC |
| Ask | Short answers grounded in a synthetic/user ledger |
| Learn | Tutor-style explanations (SIP, EMI, tax, budgeting) |
| Analytics | Spending-pattern tips without inventing numbers |
Languages: English, Hindi, Hinglish, Tamil, Telugu, Marathi, Bengali (+ more Indic mix in train data).
Key results (frozen test)
| Metric | FinnAI v4 | Qwen3-1.7B base | Qwen2.5-1.5B |
|---|---|---|---|
| SMS R-EM % (n=2028) | 94.87 | 29.39 | 49.9 |
| Amount EM % | 99.7 | 78.75 | 81.95 |
| JSON valid % | 100.0 | 100.0 | 99.65 |
| Chat groundedness % (n=78) | 80.77 | 47.44 | 65.38 |
| False-parse % | 0.25 | 100.0 | 12.88 |
Protocol: docs/llm-eval-protocol.md. Val SMS R-EM during train: 98.24%.
Quick start (Transformers)
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "finndot/finnai-slm-v4"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id, torch_dtype="bfloat16", device_map="auto"
)
messages = [
{
"role": "system",
"content": (
"Extract transaction details from this SMS as JSON: "
"{amount, merchant, type, account, balance, category}. "
"type is one of: INCOME, EXPENSE, CREDIT, TRANSFER, INVESTMENT. "
"If this is not a transaction (OTP, promo, KYC), return {}."
),
},
{
"role": "user",
"content": "HDFC Bank: Rs.499.00 debited from A/c XX1234 to SWIGGY via UPI. Avl Bal Rs.15000.50",
},
]
text = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Load LoRA-only: adapter/ in this repo (PEFT), base Qwen/Qwen3-1.7B or prior FinnAI merged.
Training details
| Parameter | Value |
|---|---|
| Base for continual train | FinnAI v3 merged (from Qwen3-1.7B) |
| Method | QLoRA SFT (4-bit NF4, bf16 compute) |
| LoRA r / α / dropout | 16 / 32 / 0.05 |
| Targets | q, k, v, o, gate, up, down_proj |
| Epochs / LR | 3 / 1.5e-4 cosine, 3% warmup |
| Effective batch | 16 (2 × 8) |
| Max seq | 1536 |
| Train size | 40,000 (SMS 20k, chat 16.8k, general 3.2k) |
| Thinking | Disabled (enable_thinking=false) |
| Hardware | 1× NVIDIA A10G (g5.2xlarge, us-east-2) |
| Seed | 42 |
Privacy: fully synthetic — no user SMS, no real accounts.
On-device (LiteRT-LM)
Production file: finndot/finnai-slm-v4-litertlm
- Format:
.litertlm, INT4dynamic_int4_block32, KV cache 1280,nothink - Size: 973,979,088 bytes
- SHA256:
36e5ec021730f0e060b78c2371d1e0011c34f28c407cdf9f993232b5cd87e7c0
Limitations
- Optimized for Indian banking SMS formats; weaker on US/EU-only banks (some hard cases remain: FAB, Schwab, etc.).
- Merchant field remains hardest; regex
parser-corestill preferred when a bank parser exists. - Chat can still paraphrase; groundedness is improved but not perfect.
- 1.7B — task-specialized, not a general assistant.
- On-device latency gate is measured separately on device hardware.
Reproduce
git clone https://github.com/AnomaProAI/Finndot
cd Finndot/ml
# dataset build + train configs: train/train_config_v4.yaml, data/out_v4_diverse/
python train/sft_qlora.py --config train/train_config_v4.yaml
python -m eval.run_eval --backend hf \
--sms data/out_v4_diverse/test_sms.jsonl \
--chat data/out_v4_diverse/chat_eval.jsonl \
--model-id train/output/v4/merged \
--baseline-id Qwen/Qwen3-1.7B \
--tag v20261005-finnai-slm-v4
Citation
@misc{finnai-slm-v4-2026,
title={FinnAI SLM v4: App-grounded on-device Indian bank SMS parsing and finance coaching},
author={FinnDot Team},
year={2026},
url={https://huggingface.co/finndot/finnai-slm-v4},
note={Apache 2.0. Continual QLoRA on Qwen3-1.7B / FinnAI v3.}
}
License
Apache 2.0 — same as Qwen3-1.7B.
- Downloads last month
- 230
Model tree for finndot/finnai-slm-v4
Base model
Qwen/Qwen3-1.7B-BaseDataset used to train finndot/finnai-slm-v4
Evaluation results
- Strict Record Exact-Match (R-EM) on FinnAI v4 SMS test (synthetic, 2028)self-reported94.870
- Amount Exact-Match on FinnAI v4 SMS test (synthetic, 2028)self-reported99.700
- JSON Validity on FinnAI v4 SMS test (synthetic, 2028)self-reported100.000
- Merchant Exact-Match on FinnAI v4 SMS test (synthetic, 2028)self-reported98.960
- Groundedness on FinnAI v4 chat_eval (78)self-reported80.770