ZaandaTeika commited on
Commit
0a13e27
·
verified ·
1 Parent(s): 6de77bf

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +23 -0
README.md CHANGED
@@ -9,6 +9,7 @@ tags:
9
  - reward model
10
  - math
11
  - hallucination-detection
 
12
  base_model:
13
  - Qwen/Qwen2.5-Math-1.5B-Instruct
14
  datasets:
@@ -120,3 +121,25 @@ token_masks = input_ids == step_sep_id
120
  step_reward = make_step_rewards(outputs[0], token_masks)
121
  print(step_reward) # one score per step
122
  ```
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9
  - reward model
10
  - math
11
  - hallucination-detection
12
+ license: apache-2.0
13
  base_model:
14
  - Qwen/Qwen2.5-Math-1.5B-Instruct
15
  datasets:
 
121
  step_reward = make_step_rewards(outputs[0], token_masks)
122
  print(step_reward) # one score per step
123
  ```
124
+
125
+ ## Source Model and Attribution
126
+
127
+ This checkpoint is a fine-tuned derivative of
128
+ [Qwen/Qwen2.5-Math-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Math-1.5B-Instruct),
129
+ released under the [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0) license.
130
+ The weights were modified: the language modelling head was replaced with a two-way
131
+ process reward head, and the model was further trained on span-derived step-level
132
+ supervision.
133
+
134
+ Training data comes from the [SHARP](https://huggingface.co/datasets/ZaandaTeika/SHARP)
135
+ corpus, whose reasoning traces and hallucination annotations are licensed under
136
+ CC BY 4.0. SHARP builds on GSM8K and MATH (both MIT); see the dataset card for the
137
+ full source attribution.
138
+
139
+ ## Additional Information
140
+
141
+ ### Licensing Information
142
+
143
+ This checkpoint is released under the
144
+ [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0) license, following its
145
+ base model.