Financial PhraseBank — Gemma 4¶
This benchmark was generated by LMTask from
trainconf/fpb-gemma4.yml on 2026-08-07T18:19:12.018977-05:00.
Results¶
Primary metric: f1 (macro) = 0.9396
Metric |
Average |
Value |
|---|---|---|
accuracy |
— |
0.9492 |
precision |
micro |
0.9492 |
recall |
micro |
0.9492 |
f1 |
micro |
0.9492 |
precision |
macro |
0.9358 |
recall |
macro |
0.9463 |
f1 |
macro |
0.9396 |
precision |
weighted |
0.9511 |
recall |
weighted |
0.9492 |
f1 |
weighted |
0.9494 |
Per-class results¶
Class |
Precision |
Recall |
F1 |
Support |
|---|---|---|---|---|
negative |
0.8684 |
0.9706 |
0.9167 |
34 |
neutral |
0.9560 |
0.9620 |
0.9590 |
158 |
positive |
0.9831 |
0.9062 |
0.9431 |
64 |
Invalid/unrecognized predictions: 0
Dataset¶
Split |
Examples |
Class counts |
|---|---|---|
train |
1752 |
|
validation |
256 |
|
test |
256 |
|
Training¶
Field |
Value |
|---|---|
Global steps |
110 |
Training loss |
18.653328254006126 |
End-to-end training time |
872.0 s |
PEFT output |
|
Adapter size |
955622176 bytes |
Trainer diagnostics¶
Metric |
Value |
|---|---|
train_runtime |
851.2863 |
train_samples_per_second |
2.058 |
train_steps_per_second |
0.129 |
total_flos |
2861561936736000.0 |
train_loss |
18.653328254006126 |
Held-out testing¶
Field |
Value |
|---|---|
Examples |
256 |
Runtime |
821.3 s |
Predictions |
|
Environment¶
Field |
Value |
|---|---|
LMTask revision |
|
Git describe |
|
Dirty working tree |
True |
Python |
3.13.13 |
OS |
Linux-4.18.0-553.124.1.el8_10.x86_64-x86_64-with-glibc2.28 |
Kernel |
4.18.0-553.124.1.el8_10.x86_64 |
CUDA_VISIBLE_DEVICES |
|
PyTorch CUDA runtime |
12.8 |
NVIDIA driver |
580.142 |
GPUs visible to the benchmark process¶
Device |
GPU |
Memory |
Peak allocated |
Peak reserved |
|---|---|---|---|---|
0 |
NVIDIA RTX A6000 |
50897289216 |
— |
— |
1 |
NVIDIA RTX A6000 |
50897289216 |
— |
— |
2 |
NVIDIA RTX A6000 |
50897289216 |
— |
— |
3 |
NVIDIA RTX A6000 |
50897289216 |
— |
— |
Package versions¶
Package |
Version |
|---|---|
scikit-learn |
1.9.0 |
torch |
2.10.0 |
transformers |
5.5.4 |
trl |
1.0.0 |
peft |
0.18.1 |
accelerate |
1.13.0 |
datasets |
4.8.5 |
pandas |
2.3.3 |
pyarrow |
24.0.0 |
huggingface-hub |
1.16.1 |
Reproducibility¶
The JSON record fpb_gemma4.json is the machine-readable source of
truth for this report. The held-out predictions are stored in
fpb_gemma4.jsonl.