Modelwire
Subscribe

Cross-Silo De-Anonymization Under Local Differential Privacy: Threat Model, Phase Transition, and Coordination Necessity

Illustration accompanying: Cross-Silo De-Anonymization Under Local Differential Privacy: Threat Model, Phase Transition, and Coordination Necessity

Researchers formalize a critical gap in federated learning privacy guarantees: standard differential privacy composition bounds assume worst-case leakage but don't predict when adversaries can actually re-identify individuals across multiple data silos. This work introduces cross-silo person-level DP and proves de-anonymization exhibits a sharp phase transition, meaning privacy degrades catastrophically once a person's records exceed a threshold number of independent sources. The finding matters for any multi-party ML system (healthcare networks, financial consortia, research collaboratives) where individuals appear in k separate datasets, each independently protected. It reframes privacy risk from theoretical composition arithmetic to empirical identifiability, forcing practitioners to rethink safe k values in real deployments.

Modelwire context

Explainer

The sharp phase transition finding is the part worth slowing down on: this isn't a gradual privacy erosion as k increases, it's a cliff. Below a threshold number of data sources, re-identification is practically infeasible; above it, it becomes nearly certain. That distinction changes how you set policy, because a 'safe enough' k today can become catastrophically unsafe with one new data-sharing agreement.

This paper lands alongside the federated learning taxonomy piece ('Beyond Weights and Gradients') published the same day, which documented how FL infrastructure is already fragmenting into domain-specific protocols with distinct privacy-utility trade-offs. That taxonomy work identified the architectural choices practitioners face; this paper now quantifies one of the failure modes those choices can trigger. Together they sketch a more complete picture of where federated deployments actually break. The synthetic data auditing paper ('Phantoms and Disclosures') is also adjacent, since both works push toward empirical identifiability as the operative privacy metric rather than theoretical composition bounds.

Watch whether healthcare or financial consortium deployments begin publishing explicit k-threshold analyses in their privacy impact assessments within the next 12 to 18 months. Adoption of that specific framing in regulatory filings would signal this formalism is crossing from academic to operational use.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsDifferential Privacy · Federated Learning · Pufferfish Privacy · Cross-Silo Person-Level DP

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Cross-Silo De-Anonymization Under Local Differential Privacy: Threat Model, Phase Transition, and Coordination Necessity · Modelwire