Instructions to use datafromlopes/llama-3.1-8b-atlassql-br with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use datafromlopes/llama-3.1-8b-atlassql-br with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.1-8B-Instruct") model = PeftModel.from_pretrained(base_model, "datafromlopes/llama-3.1-8b-atlassql-br") - Notebooks
- Google Colab
- Kaggle
llama-3.1-8b-atlassql-br
LoRA adapter of meta-llama/Llama-3.1-8B-Instruct
fine-tuned on AtlasSQL-BR, a Brazilian
Portuguese geospatial Text-to-SQL dataset over the CulturaEduca PostGIS database (schools, public
facilities and the seven-level IBGE territorial hierarchy). The model translates a question in
Portuguese into a PostGIS SQL query. No schema is injected in the prompt: the adapter is specific
to the CulturaEduca schema.
This is the model reported in the master's thesis AtlasSQL-BR: a Brazilian Portuguese geospatial Text-to-SQL dataset (IME-USP, 2026); a first version of the experiment was published at SBBD 2026.
Training
| Setting | Value |
|---|---|
| Data | thesis-split tag of the dataset: 784 training / 196 validation pairs, stratified by complexity tier, spatial function and territorial division |
| Prompt | Traduza para SQL: {question} (no schema) |
| LoRA | q_proj, v_proj; r = 8, alpha = 32, dropout = 0.1 (< 0.5 % trainable parameters) |
| Optimization | 10 epochs, batch 2 x 2 accumulation, AdamW, lr 5e-5 (linear), weight decay 0.01, 612 warmup steps, bf16, gradient checkpointing |
| Checkpoint | epoch 6, minimum validation loss (0.172; base model 1.180) |
| Hardware | Apple MacBook Pro, M5 Pro, 48 GB unified memory (PyTorch MPS); 2 h 44 min |
Results (196-pair validation split)
| Metric | Base model | This adapter |
|---|---|---|
| Execution Accuracy (%) | 0.0 | 15.8 |
| Executable Rate (%) | 0.0 | 52.6 |
| Token F1 | 0.285 | 0.757 |
| Structural F1 | 0.169 | 0.753 |
| Geospatial Function F1 | 0.108 | 0.639 |
| Spatial Exact Match (%) | 0.5 | 28.1 |
Full breakdowns (by tier, territorial division and spatial function) and the evaluation code are in the code repository.
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = "meta-llama/Llama-3.1-8B-Instruct"
tok = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype=torch.bfloat16)
model = PeftModel.from_pretrained(model, "datafromlopes/llama-3.1-8b-atlassql-br").eval()
prompt = "Traduza para SQL: Quais escolas estão a até 2 km da biblioteca municipal de Campinas?"
ids = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**ids, max_new_tokens=512, do_sample=False)
print(tok.decode(out[0][ids["input_ids"].shape[1]:], skip_special_tokens=True))
The output should be post-processed as in the evaluation code (strip code fences and annotation
tags, keep the first WITH/SELECT statement).
Citation
@inproceedings{lopes2026atlassql,
title = {AtlasSQL-BR: A Brazilian Portuguese Geospatial Text-to-SQL Dataset with Spatial Hierarchies},
author = {Lopes, Diego O. and Braghetto, Kelly R.},
booktitle = {Proceedings of the 41st Brazilian Symposium on Databases (SBBD)},
year = {2026}
}
License: MIT for the adapter weights; the base model is subject to the Llama 3.1 Community License.
- Downloads last month
- 15
Model tree for datafromlopes/llama-3.1-8b-atlassql-br
Base model
meta-llama/Llama-3.1-8B