Qwen3-4B-MegaR3ASONER-v1

A full merged reasoning model created by merging the MegaR3ASONER LoRA into Qwen/Qwen3-4B-Thinking-2507.

Model type

This repository contains the complete merged Transformers model, not only the PEFT adapter. It can be loaded directly without PeftModel.

  • Base model: Qwen/Qwen3-4B-Thinking-2507
  • Merge dtype: BF16
  • Adapter source: RexTRO111/Qwen3-4B-MegaR3ASONER-LoRA-v1
  • Training hardware: NVIDIA A10G on Modal
  • Merge hardware: NVIDIA A10G on Modal

Preliminary evaluation

On the first 100 examples selected by EleutherAI's gsm8k_cot task:

  • Flexible extraction exact match: 88%
  • Strict match: 83%

This was a limited 100-question run, not a full GSM8K score and not a controlled base-versus-fine-tune comparison.

Loading

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "RexTRO111/Qwen3-4B-MegaR3ASONER-v1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
model.eval()

Limitations

  • It may produce an incorrect intermediate thought before correcting itself.
  • It can overthink simple prompts.
  • Long reasoning traces increase latency and cost.
  • Benchmark contamination has not been exhaustively ruled out.
  • Verify answers before high-stakes use.

Licensing note

The Qwen base model and every training dataset retain their own licenses and upstream terms. Review all applicable terms before redistribution or commercial use.

Downloads last month
206
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RexTRO111/Qwen3-4B-MegaR3ASONER-v1

Finetuned
(262)
this model
Quantizations
2 models

Datasets used to train RexTRO111/Qwen3-4B-MegaR3ASONER-v1