1.75 GB
6 files
Updated 3 months ago
Name
Size
.gitattributes1.57 kB
xet
README.md1.91 kB
xet
config.json927 Bytes
xet
model.safetensors1.74 GB
xet
tokenizer.json8.34 MB
xet
tokenizer_config.json526 Bytes
xet
README.md

microsoft/deberta-v3-large Fine-tuned on PiC Phrase Similarity (Angle 3 — with_rationale)

Cross-encoder model for the PiC (Phrase in Context) phrase-similarity binary classification task. Both sentences — plus an LLM-generated rationale — are fed together so the model can attend across them simultaneously.

Base model

microsoft/deberta-v3-large

Input format

[CLS] sentence1_marked [SEP] sentence2_marked [SEP] rationale [SEP]

Target Phrase Marking

The target phrase is wrapped with <TGT>phrase</TGT> using substring matching against the raw sentence (whole-word boundaries preferred, case-insensitive fallback), since PS phrases can be multi-word and have no explicit position column in the source dataset.

Performance

Split Accuracy
Validation 0.7960
Test 0.7945

Usage

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

tokenizer = AutoTokenizer.from_pretrained("Deehan1866/ps-angle3-masked")
model = AutoModelForSequenceClassification.from_pretrained("Deehan1866/ps-angle3-masked")

s1 = "In 1990, Petit accepted a full time overnight on <TGT>air position</TGT> at gospel radio station WYLD AM."
s2 = "In 1990, Petit accepted a full time overnight on <TGT>posture while jumping</TGT> at gospel radio station WYLD AM."
rationale = "The first phrase denotes a broadcasting role; the second describes a physical stance during a jump."

sep = tokenizer.sep_token
enc = tokenizer(s1, s2 + " " + sep + " " + rationale,
                return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
    logits = model(**enc).logits
pred = torch.argmax(logits).item()
print("Similar" if pred == 1 else "Not similar")
Total size
1.75 GB
Files
6
Last updated
Jul 14
Pre-warmed CDN
US EU US EU

Contributors