Cyber-Prime 1.2 (2.6B)

Cyber-Prime 1.2 is a cybersecurity-focused instruction-tuned model in the Liquid Foundation Model (LFM2) family. It is built for defensive security analysis and cybersecurity text tasks: extracting entities from threat reports, classifying phishing emails and anomalous HTTP requests, answering cybersecurity multiple-choice questions, and producing short threat-intelligence summaries.

The checkpoint is a full, merged model. Its lineage starts from Liquid AI's LFM2-2.6B base through the Cyber-Prime series; the final release is a merged interpolation of a multi-stage checkpoint soup and a replay-stabilized Cyber-Prime branch, selected by gated evaluation on held-out validation and confirmed on the locked test split.

Benchmark results

The table uses the same models and scores shown in the benchmark graphic. Values are rounded to three decimals.

Benchmark Cyber-Prime 1 Cyber-Prime 1.1 Cyber-Prime 1.2 GPT-4 GPT-3.5 Turbo Mistral-7B-Instruct Llama-2-7B
CyNER (F1) 0.382 0.442 0.687 0.554 0.334 0.323 0.263
APTNER (F1) 0.413 0.446 0.507 0.500 0.409 0.262 0.280
CyNews (ROUGE-1) 0.354 0.441 0.461 0.275 0.271 0.217 0.003
SecMMLU (Acc) 0.580 0.600 0.700 0.830 0.780 0.720 0.630
CyQuiz (Acc) 0.570 0.600 0.750 0.810 0.830 0.690 0.620
Email Phishing Detection (F1) 0.728 0.888 0.979 0.939 0.789 0.889 0.942
HTTP Attack Log Analysis (F1) 0.483 0.571 0.890 0.841 0.831 0.472 0.428
Overall Cybersecurity Average 0.501 0.570 0.711 0.721 0.609 0.511 0.451

Overall average is the unweighted mean of the seven benchmark scores shown; it is a descriptive comparison, not an official CyberBench aggregate.

Cyber-Prime 1.2 compared with earlier Cyber-Prime releases and published CyberBench baselines

Evaluation protocol

  • Test data were held out from all training mixes; evaluation used the locked CyberBench test split.
  • CyNER, APTNER, SecMMLU, CyQuiz, email, and HTTP were evaluated with five retrieval-selected in-context examples from their training pools (the CyberBench paper protocol). CyNews was evaluated zero-shot.
  • Generation used greedy decoding (temperature 0) with the model's chat template; reasoning text was excluded before task scoring where applicable.
  • Email F1 treats phishing as the positive class; HTTP F1 treats anomalous as the positive class.
  • Cyber-Prime 1.0 and the external baselines are reproduced from the published CyberBench results and the prior Cyber-Prime release graphics; they were not re-run here. The Cyber-Prime 1.1 column is a prior-run reference under the same harness, which differs from the published 1.1 card values. Cross-model bars are useful context, not a strictly matched comparison.

Training

Cyber-Prime 1.2 was post-trained on curated, license-checked cybersecurity data with strict benchmark decontamination, combining supervised fine-tuning, verified reasoning traces, and model merging. The released weights are a linear weight interpolation between two internal checkpoints — a multi-stage Cyber-Prime soup and a replay-stabilized branch — selected by gated evaluation rather than train metrics.

Intended use

  • Defensive cybersecurity education and research.
  • Cybersecurity entity extraction from reports and logs.
  • Triage assistance for phishing-email and anomalous-HTTP classification.
  • Cybersecurity multiple-choice question answering.
  • Short threat-intelligence headline and summary generation.

Use the model as an assistive component with human review. It is not a substitute for security controls, incident-response procedures, or expert validation, and it is not a reliable autonomous vulnerability assessment or exploitation agent.

Limitations

  • Performance is measured on the listed benchmark test sets; it should not be assumed to transfer to every organization, threat actor, protocol, or language.
  • NER recall remains lower than precision, especially for rare or densely packed entities.
  • The MCQ test sets are small (100 examples each), so small score differences may be noise.
  • The baseline comparison mixes different evaluation runs and harnesses; prompts and serving stacks can materially change scores.
  • The model can produce incorrect, outdated, or overconfident cybersecurity claims. Verify indicators, classifications, and recommendations independently before taking action.

Loading

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "Akahsizrr/Cyber-Prime-1.2"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
    repo,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

messages = [{
    "role": "user",
    "content": "Classify this HTTP request as normal or anomalous: GET /index.html HTTP/1.1",
}]
input_ids = tokenizer.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_tensors="pt",
).to(model.device)
output = model.generate(input_ids, max_new_tokens=256, do_sample=False)
print(tokenizer.decode(output[0][input_ids.shape[-1]:], skip_special_tokens=True))

Use task-specific instructions and validate the output format expected by your downstream system. For NER, request a JSON object; for email and HTTP classification, request only the benchmark label; for CyNews-like summarization, request a concise headline.

License

This model is derived from LiquidAI/LFM2-2.6B and is distributed under the upstream LFM Open License v1.0, included in LICENSE. This is a custom license, not Apache-2.0. In particular, Section 5 does not license commercial use by a legal entity with annual revenue of USD 10 million or more; such use requires separate permission from the licensor. Review the complete license before use or redistribution.

Data and references

Downloads last month
25
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Akahsizrr/Cyber-Prime-1.2

Finetuned
(44)
this model
Quantizations
1 model

Datasets used to train Akahsizrr/Cyber-Prime-1.2

Collection including Akahsizrr/Cyber-Prime-1.2