Abstract
Financial text, textbooks, and question–answer pairs are abundant, but only a small fraction is directly usable for reasoning-focused post-training. We present a data-centric pipeline that constructs complementary corpora by mining open-source reasoning traces, distilling financial instruction data, and generating knowledge-graph-guided question–answer pairs from financial educational material. Three lightweight classifiers select finance-relevant examples, reject under-specified questions, and identify tasks suitable for reinforcement learning with compact rule-based verifiers. We evaluate supervised fine-tuning, self-distilled fine-tuning, model merging, and GRPO on FINESSE-Bench.