ArticleScientific reports2026
Use large language model to enhance reasoning of another large language model through reward updated GRPO.
Yiqiao YinPubMed ↗Full text ↗Publisher ↗
No numbers read from the abstract.
ArticleScientific reports2026
Yiqiao YinPubMed ↗Full text ↗Publisher ↗
No numbers read from the abstract.