Modelwire
Subscribe

Audio foundation models recover evolutionary relationships without training for it

Foundation audio models trained on human speech and environmental sound encode evolutionary structure without explicit supervision, as demonstrated by their ability to recover phylogenetic relationships across marine mammal species. This finding reveals that large pretrained models capture biological organization at scales far beyond their training objectives, raising questions about what latent structure emerges in high-dimensional embeddings. The result matters for model interpretability and suggests that foundation models may be learning compressed representations of natural hierarchies, with implications for how we evaluate and deploy these systems across domains.

Modelwire context

Explainer

The real finding is negative: domain-specific audio pretraining (BEATs-bio) underperformed generic models (CLAP, AST) at recovering evolutionary structure. This inverts the usual intuition that specialized training data improves performance on specialized tasks.

This connects directly to the interpretability thread from the bag-of-waves EEG work published the same day. Both papers ask what structure emerges when you don't explicitly supervise for it: one shows unsupervised waveform templates reveal clinically meaningful patterns, the other shows unsupervised embeddings recover phylogenetic hierarchies. Together they suggest that large models trained on generic objectives may already capture domain structure implicitly, raising a question about whether domain-specific pretraining is actually necessary or if it introduces noise that obscures latent organization.

If researchers apply the same phylogenetic recovery test to BEATs-bio's learned representations versus its outputs (testing whether the structure is in the model but masked by the domain-specific training objective), that would confirm whether specialization is actively harmful or merely redundant. If the structure disappears entirely on bird vocalizations or terrestrial mammals, the finding is narrower than claimed.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsAST · CLAP · BEATs-bio · BirdNET · Watkins Marine Mammal Sound Database

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.LG originally reported this story as Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Audio foundation models recover evolutionary relationships without training for it · Modelwire