← Evidence map

ArticleScientific reports2026

Use large language model to enhance reasoning of another large language model through reward updated GRPO.

Yiqiao YinPubMed ↗Full text ↗Publisher ↗

No numbers read from the abstract.

Not cited yet

Full record →Abstract, authors, funding and every citing paper · PMID 41673205