Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
Large Language Models (LLMs) have been playing a transformative role in natural language understanding and generation, yet adapting LLMs to domain-specific and privacy-sensitive data remains challenging under centralized training. Federated Learning (FL) provides a promising alternative by enabling training LLMs collaboratively without sharing raw data. However, integrating FL and LLMs introduces new challenges, including model size, device heterogeneity, non-IID data, and alignment requirements. This survey offers a structured overview of the federated LLM ecosystem. We present a comprehensive taxonomy encompassing system architectures, advanced data strategies for addressing heterogeneity, and retrieval-augmented generation in federated contexts. Additionally, we review efficient adaptation methods that enable LLM tuning on resource-constrained clients and analyze data security and privacy concerns. We conclude by summarizing emerging applications in healthcare, industry, software engineering, and finance, and by outlining open problems and research opportunities for scalable, secure, and responsible federated LLM deployment.
This work is licensed under a Creative Commons Attribution 4.0 International License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Comments on this article