Staged Continual Adaptation of Multimodal Foundation Models for Japanese Financial Documents
Genshin Kakimoto ⋅ Atsushi Yanagisawa
Abstract
Staged post-training of multimodal foundation models pairs vision--language alignment, reasoning distillation, and domain tuning, but how each phase \emph{trades} capabilities against the others is largely under-characterised. Tracking an 8B multimodal model from an un-fine-tuned baseline through three training phases on Japanese financial disclosures, we find that each benchmark peaks at a \emph{different} phase, so the best checkpoint is task-dependent---with implications for checkpoint retention and compute budgeting in resource-constrained continual adaptation.
Chat is not available.
Successful Page Load