Biologically-inspired vision transformer gains robustness through selective token processing
Researchers have developed Foveated Dynamic Transformer (FDT), a vision transformer that mimics biological foveation to achieve both computational efficiency and robustness against noise and adversarial attacks without explicit adversarial training. By selectively processing high-priority image regions through fixation and foveation modules, FDT reduces token overhead while gaining emergent resilience properties. This biologically-inspired approach addresses a persistent tension in vision model design: efficiency gains typically come at the cost of robustness. The work signals growing interest in architectural inductive biases drawn from neuroscience as a path to more capable and efficient models, potentially influencing how future vision systems balance speed, accuracy, and adversarial resilience.62























