# Financial PhraseBank — Gemma 4 This benchmark was generated by LMTask from `trainconf/fpb-gemma4.yml` on 2026-08-07T18:19:12.018977-05:00. ## Results **Primary metric:** f1 (macro) = **0.9396** | Metric | Average | Value | |---|---|---:| | accuracy | — | 0.9492 | | precision | micro | 0.9492 | | recall | micro | 0.9492 | | f1 | micro | 0.9492 | | precision | macro | 0.9358 | | recall | macro | 0.9463 | | f1 | macro | 0.9396 | | precision | weighted | 0.9511 | | recall | weighted | 0.9492 | | f1 | weighted | 0.9494 | ### Per-class results | Class | Precision | Recall | F1 | Support | |---|---:|---:|---:|---:| | negative | 0.8684 | 0.9706 | 0.9167 | 34 | | neutral | 0.9560 | 0.9620 | 0.9590 | 158 | | positive | 0.9831 | 0.9062 | 0.9431 | 64 | Invalid/unrecognized predictions: **0** ## Dataset | Split | Examples | Class counts | |---|---:|---| | train | 1752 | `positive`=441, `neutral`=1076, `negative`=235 | | validation | 256 | `neutral`=157, `negative`=34, `positive`=65 | | test | 256 | `neutral`=158, `negative`=34, `positive`=64 | ## Training | Field | Value | |---|---| | Global steps | 110 | | Training loss | 18.653328254006126 | | End-to-end training time | 872.0 s | | PEFT output | `data/fpb/gemma4/model/peft` | | Adapter size | 955622176 bytes | ### Trainer diagnostics | Metric | Value | |---|---:| | train_runtime | 851.2863 | | train_samples_per_second | 2.058 | | train_steps_per_second | 0.129 | | total_flos | 2861561936736000.0 | | train_loss | 18.653328254006126 | ## Held-out testing | Field | Value | |---|---| | Examples | 256 | | Runtime | 821.3 s | | Predictions | `data/fpb/gemma4/benchmark/fpb_gemma4.jsonl` | ## Environment | Field | Value | |---|---| | LMTask revision | `e4deeee630568375c248f6a597f3e5137e06e0c4` | | Git describe | `v0.0.1-101-ge4deeee-dirty` | | Dirty working tree | True | | Python | 3.13.13 | | OS | Linux-4.18.0-553.124.1.el8_10.x86_64-x86_64-with-glibc2.28 | | Kernel | 4.18.0-553.124.1.el8_10.x86_64 | | CUDA_VISIBLE_DEVICES | `4,5,6,7` | | PyTorch CUDA runtime | 12.8 | | NVIDIA driver | 580.142 | ### GPUs visible to the benchmark process | Device | GPU | Memory | Peak allocated | Peak reserved | |---:|---|---:|---:|---:| | 0 | NVIDIA RTX A6000 | 50897289216 | — | — | | 1 | NVIDIA RTX A6000 | 50897289216 | — | — | | 2 | NVIDIA RTX A6000 | 50897289216 | — | — | | 3 | NVIDIA RTX A6000 | 50897289216 | — | — | ### Package versions | Package | Version | |---|---| | scikit-learn | 1.9.0 | | torch | 2.10.0 | | transformers | 5.5.4 | | trl | 1.0.0 | | peft | 0.18.1 | | accelerate | 1.13.0 | | datasets | 4.8.5 | | pandas | 2.3.3 | | pyarrow | 24.0.0 | | huggingface-hub | 1.16.1 | ## Reproducibility The JSON record `fpb_gemma4.json` is the machine-readable source of truth for this report. The held-out predictions are stored in `fpb_gemma4.jsonl`.