Natural language emerges as primary feedback channel for agent training
Researchers have formalized Verbal Reinforcement Learning, a framework where natural language serves as the primary feedback mechanism for training language agents. Rather than relying solely on numerical rewards or parameter updates, VRL leverages human-interpretable text to convey task definitions, real-time reasoning guidance, and learning signals. The taxonomy identifies three distinct applications: language as task specification, language as in-context steering during inference, and language as a training signal. This shift matters because it bridges human intent and model optimization in ways that scale with LLM capabilities, potentially reducing the need for expensive labeled datasets while improving alignment between agent behavior and human preferences.62