Researchers propose explicit concept design as core LLM architecture choice
Researchers propose treating concept structure as an explicit design choice for LLMs rather than a post-hoc discovery. Current models encode concepts implicitly through statistical patterns, forcing interpretability work to reverse-engineer meaning after training. This paper maps a design space across pipeline stages (training, architecture, inference, interpretation) and internal versus external concept representation, arguing that deliberate architectural choices could yield models with stronger compositionality, controllability, and human-aligned reasoning. The shift from found to designed concepts addresses a core limitation in model transparency and could reshape how future architectures balance capability with interpretability.62




















