The Scissors Effect: When Resize-Based Input Diversity Helps or Hurts Transfer Attacks

A new study reveals that input diversity, a standard technique in adversarial attack research, produces opposite effects depending on the target model's training regime. When attacking robustly trained models, randomized resizing and padding actually degrades transfer success by up to 10.3% on ImageNet, contradicting the field's conventional wisdom. This finding reshapes how researchers should calibrate attack strategies across different model architectures (CNNs, ViTs, Swin, ConvNeXt) and exposes a critical blind spot in transfer-based adversarial robustness evaluation that has persisted across a decade of attack methods.
Modelwire context
ExplainerThe finding isn't just that input diversity sometimes fails; it's that the field has been evaluating adversarial robustness using attack pipelines that systematically flatter robust models, meaning published robustness benchmarks may be artificially inflated wherever input diversity was applied against robustly trained targets.
This connects directly to the broader problem of evaluation reliability that has surfaced repeatedly in recent coverage. The 'Robust Diffusion Models via Divergence-Induced Weighted Denoising' paper from the same day addresses a related gap: that standard training objectives don't hold up under realistic data conditions, and that practitioners often discover this only after deployment. Both papers are pointing at the same structural issue from different angles, namely that the assumptions baked into standard pipelines (loss functions, attack augmentations) quietly degrade reliability in ways that aren't visible until someone runs the controlled ablation. The scissors effect described here is essentially an evaluation contamination problem, not just an attack design problem.
Watch whether benchmark suites like RobustBench update their leaderboard methodology to control for input diversity settings across robust versus standard models. If they do within the next two release cycles, this finding has been accepted as a validity concern rather than a niche result.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsImageNet · CIFAR-10 · CNN · ViT · Swin · ConvNeXt
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.