Financial PhraseBank — DeepSeek-R1-Distill-Qwen 3¶
This benchmark was generated by LMTask from
trainconf/fpb-dsr1qwen3.yml on 2026-08-08T16:03:28.602933-05:00.
Results¶
Primary metric: f1 (macro) = 0.8147
Metric |
Average |
Value |
|---|---|---|
accuracy |
— |
0.7969 |
precision |
micro |
0.7969 |
recall |
micro |
0.7969 |
f1 |
micro |
0.7969 |
precision |
macro |
0.8068 |
recall |
macro |
0.8687 |
f1 |
macro |
0.8147 |
precision |
weighted |
0.8580 |
recall |
weighted |
0.7969 |
f1 |
weighted |
0.8023 |
Per-class results¶
Class |
Precision |
Recall |
F1 |
Support |
|---|---|---|---|---|
negative |
0.8649 |
0.9412 |
0.9014 |
34 |
neutral |
0.9649 |
0.6962 |
0.8088 |
158 |
positive |
0.5905 |
0.9688 |
0.7337 |
64 |
Invalid/unrecognized predictions: 0
Dataset¶
Split |
Examples |
Class counts |
|---|---|---|
train |
1752 |
|
validation |
256 |
|
test |
256 |
|
Training¶
Field |
Value |
|---|---|
Global steps |
110 |
Training loss |
2.509641246362166 |
End-to-end training time |
324.9 s |
PEFT output |
|
Adapter size |
1291899160 bytes |
Trainer diagnostics¶
Metric |
Value |
|---|---|
train_runtime |
312.8679 |
train_samples_per_second |
5.6 |
train_steps_per_second |
0.352 |
total_flos |
7265346283044864.0 |
train_loss |
2.509641246362166 |
Held-out testing¶
Field |
Value |
|---|---|
Examples |
256 |
Runtime |
3318.5 s |
Predictions |
|
Environment¶
Field |
Value |
|---|---|
LMTask revision |
|
Git describe |
|
Dirty working tree |
True |
Python |
3.13.13 |
OS |
Linux-4.18.0-553.124.1.el8_10.x86_64-x86_64-with-glibc2.28 |
Kernel |
4.18.0-553.124.1.el8_10.x86_64 |
CUDA_VISIBLE_DEVICES |
|
PyTorch CUDA runtime |
12.8 |
NVIDIA driver |
580.142 |
GPUs visible to the benchmark process¶
Device |
GPU |
Memory |
Peak allocated |
Peak reserved |
|---|---|---|---|---|
0 |
NVIDIA RTX A6000 |
50897289216 |
— |
— |
1 |
NVIDIA RTX A6000 |
50897289216 |
— |
— |
2 |
NVIDIA RTX A6000 |
50897289216 |
— |
— |
3 |
NVIDIA RTX A6000 |
50897289216 |
— |
— |
Package versions¶
Package |
Version |
|---|---|
scikit-learn |
1.9.0 |
torch |
2.10.0 |
transformers |
5.5.4 |
trl |
1.0.0 |
peft |
0.18.1 |
accelerate |
1.13.0 |
datasets |
4.8.5 |
pandas |
2.3.3 |
pyarrow |
24.0.0 |
huggingface-hub |
1.16.1 |
Reproducibility¶
The JSON record fpb_dsr1qwen3.json is the machine-readable source of
truth for this report. The held-out predictions are stored in
fpb_dsr1qwen3.jsonl.