Technical Report · 2026

Data-Centric Post-Training for Financial Reasoning

Mining, Distillation, and Verifiable Learning

Zhirayr Hayrapetyan · Andrei Kalmykov · Denis Kokosinskii · Dmitry Stanishevskii · Dmitry Zmitrovich

Abstract

Financial text, textbooks, and question–answer pairs are abundant, but only a small fraction is directly usable for reasoning-focused post-training. We present a data-centric pipeline that constructs complementary corpora by mining open-source reasoning traces, distilling financial instruction data, and generating knowledge-graph-guided question–answer pairs from financial educational material. Three lightweight classifiers select finance-relevant examples, reject under-specified questions, and identify tasks suitable for reinforcement learning with compact rule-based verifiers. We evaluate supervised fine-tuning, self-distilled fine-tuning, model merging, and GRPO on FINESSE-Bench.

Selected result

+1.0 to +2.8 pp

Self-distilled SFT improves FINESSE-Bench accuracy over the corresponding starting models, whereas ordinary SFT loses 3.2–4.0 points in the selected comparisons.

Open artifacts

Datasets and auxiliary classifiers released with this work.

Citation

arXiv:2609.10113 [cs.CL]

@article{hayrapetyan2026datacentric,
  title   = {Data-Centric Post-Training for Financial Reasoning:
             Mining, Distillation, and Verifiable Learning},
  author  = {Hayrapetyan, Zhirayr and Kalmykov, Andrei and
             Kokosinskii, Denis and Stanishevskii, Dmitry and
             Zmitrovich, Dmitry},
  journal = {arXiv preprint arXiv:2609.10113},
  year    = {2026}
}