# Financial PhraseBank — DeepSeek-R1-Distill-Qwen 3 This benchmark was generated by LMTask from `trainconf/fpb-dsr1qwen3.yml` on 2026-08-08T16:03:28.602933-05:00. ## Results **Primary metric:** f1 (macro) = **0.8147** | Metric | Average | Value | |-----------|----------|-------:| | accuracy | — | 0.7969 | | precision | micro | 0.7969 | | recall | micro | 0.7969 | | f1 | micro | 0.7969 | | precision | macro | 0.8068 | | recall | macro | 0.8687 | | f1 | macro | 0.8147 | | precision | weighted | 0.8580 | | recall | weighted | 0.7969 | | f1 | weighted | 0.8023 | ### Per-class results | Class | Precision | Recall | F1 | Support | |----------|----------:|-------:|-------:|--------:| | negative | 0.8649 | 0.9412 | 0.9014 | 34 | | neutral | 0.9649 | 0.6962 | 0.8088 | 158 | | positive | 0.5905 | 0.9688 | 0.7337 | 64 | Invalid/unrecognized predictions: **0** ## Dataset | Split | Examples | Class counts | |------------|---------:|------------------------------------------------| | train | 1752 | `positive`=441, `neutral`=1076, `negative`=235 | | validation | 256 | `neutral`=157, `negative`=34, `positive`=65 | | test | 256 | `neutral`=158, `negative`=34, `positive`=64 | ## Training | Field | Value | |--------------------------|---------------------------------| | Global steps | 110 | | Training loss | 2.509641246362166 | | End-to-end training time | 324.9 s | | PEFT output | `data/fpb/dsr1qwen3/model/peft` | | Adapter size | 1291899160 bytes | ### Trainer diagnostics | Metric | Value | |--------------------------|-------------------:| | train_runtime | 312.8679 | | train_samples_per_second | 5.6 | | train_steps_per_second | 0.352 | | total_flos | 7265346283044864.0 | | train_loss | 2.509641246362166 | ## Held-out testing | Field | Value | |-------------|----------------------------------------------------| | Examples | 256 | | Runtime | 3318.5 s | | Predictions | `data/fpb/dsr1qwen3/benchmark/fpb_dsr1qwen3.jsonl` | ## Environment | Field | Value | |----------------------|------------------------------------------------------------| | LMTask revision | `51cf677d65e9b7c9deffa2878c505dba82b59d89` | | Git describe | `v0.0.1-105-g51cf677` | | Dirty working tree | True | | Python | 3.13.13 | | OS | Linux-4.18.0-553.124.1.el8_10.x86_64-x86_64-with-glibc2.28 | | Kernel | 4.18.0-553.124.1.el8_10.x86_64 | | CUDA_VISIBLE_DEVICES | `4,5,6,7` | | PyTorch CUDA runtime | 12.8 | | NVIDIA driver | 580.142 | ### GPUs visible to the benchmark process | Device | GPU | Memory | Peak allocated | Peak reserved | |-------:|------------------|------------:|---------------:|--------------:| | 0 | NVIDIA RTX A6000 | 50897289216 | — | — | | 1 | NVIDIA RTX A6000 | 50897289216 | — | — | | 2 | NVIDIA RTX A6000 | 50897289216 | — | — | | 3 | NVIDIA RTX A6000 | 50897289216 | — | — | ### Package versions | Package | Version | |-----------------|---------| | scikit-learn | 1.9.0 | | torch | 2.10.0 | | transformers | 5.5.4 | | trl | 1.0.0 | | peft | 0.18.1 | | accelerate | 1.13.0 | | datasets | 4.8.5 | | pandas | 2.3.3 | | pyarrow | 24.0.0 | | huggingface-hub | 1.16.1 | ## Reproducibility The JSON record `fpb_dsr1qwen3.json` is the machine-readable source of truth for this report. The held-out predictions are stored in `fpb_dsr1qwen3.jsonl`.