Qwen-based models tested for biomedical text generation with four alignment methods
Researchers benchmarked four post-training alignment techniques (SFT, DPO, ORPO, GRPO) on Qwen-based small language models for biomedical data-to-text tasks, specifically medication leaflet generation. The work bridges a gap in specialized domain adaptation by testing whether alignment methods designed for general-purpose LLMs transfer effectively to constrained, high-stakes biomedical outputs. Cross-dataset evaluation using FDA drug labels and dual metrics (lexical and semantic) provides practical guidance for practitioners deploying SLMs in regulated healthcare contexts where accuracy and clarity directly impact patient safety.52




























