Financial PhraseBank — Gemma 4

This benchmark was generated by LMTask from trainconf/fpb-gemma4.yml on 2026-08-07T18:19:12.018977-05:00.

Results

Primary metric: f1 (macro) = 0.9396

Metric

Average

Value

accuracy

0.9492

precision

micro

0.9492

recall

micro

0.9492

f1

micro

0.9492

precision

macro

0.9358

recall

macro

0.9463

f1

macro

0.9396

precision

weighted

0.9511

recall

weighted

0.9492

f1

weighted

0.9494

Per-class results

Class

Precision

Recall

F1

Support

negative

0.8684

0.9706

0.9167

34

neutral

0.9560

0.9620

0.9590

158

positive

0.9831

0.9062

0.9431

64

Invalid/unrecognized predictions: 0

Dataset

Split

Examples

Class counts

train

1752

positive=441, neutral=1076, negative=235

validation

256

neutral=157, negative=34, positive=65

test

256

neutral=158, negative=34, positive=64

Training

Field

Value

Global steps

110

Training loss

18.653328254006126

End-to-end training time

872.0 s

PEFT output

data/fpb/gemma4/model/peft

Adapter size

955622176 bytes

Trainer diagnostics

Metric

Value

train_runtime

851.2863

train_samples_per_second

2.058

train_steps_per_second

0.129

total_flos

2861561936736000.0

train_loss

18.653328254006126

Held-out testing

Field

Value

Examples

256

Runtime

821.3 s

Predictions

data/fpb/gemma4/benchmark/fpb_gemma4.jsonl

Environment

Field

Value

LMTask revision

e4deeee630568375c248f6a597f3e5137e06e0c4

Git describe

v0.0.1-101-ge4deeee-dirty

Dirty working tree

True

Python

3.13.13

OS

Linux-4.18.0-553.124.1.el8_10.x86_64-x86_64-with-glibc2.28

Kernel

4.18.0-553.124.1.el8_10.x86_64

CUDA_VISIBLE_DEVICES

4,5,6,7

PyTorch CUDA runtime

12.8

NVIDIA driver

580.142

GPUs visible to the benchmark process

Device

GPU

Memory

Peak allocated

Peak reserved

0

NVIDIA RTX A6000

50897289216

1

NVIDIA RTX A6000

50897289216

2

NVIDIA RTX A6000

50897289216

3

NVIDIA RTX A6000

50897289216

Package versions

Package

Version

scikit-learn

1.9.0

torch

2.10.0

transformers

5.5.4

trl

1.0.0

peft

0.18.1

accelerate

1.13.0

datasets

4.8.5

pandas

2.3.3

pyarrow

24.0.0

huggingface-hub

1.16.1

Reproducibility

The JSON record fpb_gemma4.json is the machine-readable source of truth for this report. The held-out predictions are stored in fpb_gemma4.jsonl.