Financial PhraseBank — DeepSeek-R1-Distill-Qwen 3

This benchmark was generated by LMTask from trainconf/fpb-dsr1qwen3.yml on 2026-08-08T16:03:28.602933-05:00.

Results

Primary metric: f1 (macro) = 0.8147

Metric

Average

Value

accuracy

0.7969

precision

micro

0.7969

recall

micro

0.7969

f1

micro

0.7969

precision

macro

0.8068

recall

macro

0.8687

f1

macro

0.8147

precision

weighted

0.8580

recall

weighted

0.7969

f1

weighted

0.8023

Per-class results

Class

Precision

Recall

F1

Support

negative

0.8649

0.9412

0.9014

34

neutral

0.9649

0.6962

0.8088

158

positive

0.5905

0.9688

0.7337

64

Invalid/unrecognized predictions: 0

Dataset

Split

Examples

Class counts

train

1752

positive=441, neutral=1076, negative=235

validation

256

neutral=157, negative=34, positive=65

test

256

neutral=158, negative=34, positive=64

Training

Field

Value

Global steps

110

Training loss

2.509641246362166

End-to-end training time

324.9 s

PEFT output

data/fpb/dsr1qwen3/model/peft

Adapter size

1291899160 bytes

Trainer diagnostics

Metric

Value

train_runtime

312.8679

train_samples_per_second

5.6

train_steps_per_second

0.352

total_flos

7265346283044864.0

train_loss

2.509641246362166

Held-out testing

Field

Value

Examples

256

Runtime

3318.5 s

Predictions

data/fpb/dsr1qwen3/benchmark/fpb_dsr1qwen3.jsonl

Environment

Field

Value

LMTask revision

51cf677d65e9b7c9deffa2878c505dba82b59d89

Git describe

v0.0.1-105-g51cf677

Dirty working tree

True

Python

3.13.13

OS

Linux-4.18.0-553.124.1.el8_10.x86_64-x86_64-with-glibc2.28

Kernel

4.18.0-553.124.1.el8_10.x86_64

CUDA_VISIBLE_DEVICES

4,5,6,7

PyTorch CUDA runtime

12.8

NVIDIA driver

580.142

GPUs visible to the benchmark process

Device

GPU

Memory

Peak allocated

Peak reserved

0

NVIDIA RTX A6000

50897289216

1

NVIDIA RTX A6000

50897289216

2

NVIDIA RTX A6000

50897289216

3

NVIDIA RTX A6000

50897289216

Package versions

Package

Version

scikit-learn

1.9.0

torch

2.10.0

transformers

5.5.4

trl

1.0.0

peft

0.18.1

accelerate

1.13.0

datasets

4.8.5

pandas

2.3.3

pyarrow

24.0.0

huggingface-hub

1.16.1

Reproducibility

The JSON record fpb_dsr1qwen3.json is the machine-readable source of truth for this report. The held-out predictions are stored in fpb_dsr1qwen3.jsonl.