UniFLoW: Universal Multi-Modal Federated LoRA Fine-Tuning Framework with Analytical Aggregation
Abstract
Lay Summary
Large AI models are often adapted to new tasks using data from many users or organizations. However, in many real-world settings, this data cannot be collected in one place because of privacy, ownership, or communication constraints. This makes it important to train models collaboratively while keeping data local. Our paper studies how to make this collaborative training more reliable and efficient when only a small part of the model is updated. Existing methods can become unstable because different participants may update the model in conflicting directions, leading to slow progress or inconsistent results. We propose a new approach that better combines the updates from different participants so that training is more stable and the final model performs better. Across language and multimodal tasks, our method improves accuracy and produces more reliable answers compared with existing approaches. The results suggest that large AI models can be adapted more effectively in privacy-sensitive distributed environments, which may be useful for applications such as healthcare, education, and personalized assistants where data should remain under local control.