Abstract
Geo-distributed training and Federated Learning (FL) provide viable solutions to address the substantial data and computational resource needs associated with training large language models (LLMs). However, we empirically demonstrate that a single attacker can significantly compromise the safety alignment of LLMs through malicious training, and existing defenses like robust aggregation or trust-based frameworks fail under this setting due to data heterogeneity. We identify two existing server-side defense strategies that effectively counter naive jailbreak attacks: Task Performance Check (TPC), which filters out model updates with low downstream performance, and Malicious Output Scrutiny (MOS), which detects harmful outputs by prompting uploaded models with malicious queries. To evade both defenses, we design a trigger-based jailbreak variant that preserves downstream performance using a novel regularization method to limit the excessive model updates on jailbreak datasets. We further conceal malicious triggers by mixing the malicious dataset with pseudo-contrastive safety-aligned answers to maintain the original safety alignment. Experiments on several widely used safety-aligned LLMs show that CloudGhost can consistently implant triggers into the global model without degrading downstream performance, achieving 74–93% attack success rate (ASR) and below 5% detection true rate (DTR).
| Original language | English |
|---|---|
| Title of host publication | The Fourteenth International Conference on Learning Representations, ICLR 2026 |
| Publisher | International Conference on Learning Representations, ICLR |
| Number of pages | 31 |
| Publication status | Published - 23 Apr 2026 |
| Event | 14th International Conference on Learning Representations, ICLR 2026 - Rio de Janeiro, Brazil Duration: 23 Apr 2026 → 27 Apr 2026 https://iclr.cc/Conferences/2026 (Conference website) https://openreview.net/group?id=ICLR.cc/2026 (Conference proceedings) https://iclr.cc/virtual/2026/calendar (Conference schedule) |
Publication series
| Name | International Conference on Learning Representations |
|---|---|
| Publisher | International Conference on Learning Representations, ICLR |
Conference
| Conference | 14th International Conference on Learning Representations, ICLR 2026 |
|---|---|
| Abbreviated title | ICLR 2026 |
| Country/Territory | Brazil |
| City | Rio de Janeiro |
| Period | 23/04/26 → 27/04/26 |
| Internet address |
|
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 9 Industry, Innovation, and Infrastructure
User-Defined Keywords
- Jailbreak attack
- Geo-distributed LLM Training
- Federated Learning
- Large Language Models
Fingerprint
Dive into the research topics of 'Ghost in the Cloud: Your Geo-Distributed Large Language Models Training is Easily Manipulated'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver