Evidence map›Paper›PMID 42495013›Full record

ArticleResearch (Washington, D.C.)2026

Deep Reinforcement Learning for Real-World Humanoid Robot Locomotion Control with Automatic Reward Learning.

Renzhi Lu, Jie Wang, Zonghe Shao, Ruijuan Chen, Lijun Zhu, Yuzhi Jiang, Yunyi Pang, Dongfang Liang, Yang Shi, Han Ding

Abstract read
In one paragraph

Article in Research (Washington, D.C.), 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Renzhi LuSchool of Artificial Intelligence and Automation, Key Laboratory of Image Processing and Intelligent Control, Engineering Research Center of Autonomous Intelligent Unmanned Systems, Chinese Ministry of Education, Huazhong University of Science and Technology, Wuhan 430074, China.
Jie WangSchool of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China.ORCID https://orcid.org/0009-0005-7337-5692
Zonghe ShaoSchool of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China.
Ruijuan ChenResearch Center for Applied Mathematics and Interdisciplinary Science, School of Mathematics and Statistics, Wuhan Textile University, Wuhan 430200, Hubei, China.
Lijun ZhuSchool of Artificial Intelligence and Automation, Key Laboratory of Image Processing and Intelligent Control, Engineering Research Center of Autonomous Intelligent Unmanned Systems, Chinese Ministry of Education, Huazhong University of Science and Technology, Wuhan 430074, China.
Yuzhi JiangSchool of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China.
Yunyi PangSchool of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China.
Dongfang LiangDepartment of Engineering, University of Cambridge, Cambridge CB2 1PZ, UK.
Yang ShiDepartment of Mechanical Engineering, University of Victoria, Victoria, BC, V8W 3P6, Canada.
Han DingState Key Laboratory of Digital Manufacturing Equipment and Technology, Huazhong University of Science and Technology, Wuhan 430074, China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Humanoid robots possess the potential to solve complex problems across diverse environments, such as nuclear-contaminated zones, epidemic-affected areas, and extraterrestrial missions. However, a humanoid robot is an inherently complex system that integrates multiple disciplines, including perception, mechanical design, materials science, and motion control, each of which requires comprehensive and in-depth investigation. Among these aspects, motion control plays a crucial role, as it directly determines the robot's motion accuracy, stability, and flexibility. In recent years, with the rapid evolution of graphics-processing-unit-based parallel computing and high-fidelity simulation environments, various deep-reinforcement-learning (DRL)-based approaches have been proposed to achieve precise and robust motion control due to its flexibility and adaptability in uncertain and dynamic environments. However, the inherent complexity and uncertainty of real-world tasks pose substantial challenges when designing effective reward functions for DRL agents. Most current methods typically rely on manually engineered or externally tuned reward signals and therefore require considerable domain expertise, associated with considerable human efforts and a long convergence time; these issues may even trigger mission failure. This work proposes an automatic reward learning method to derive reward functions for DRL in humanoid robot locomotion control. Specifically, a bilevel optimization framework is developed to enable automatic reward learning during policy learning. The reward learning mechanism in the upper level adaptively constructs and optimizes the reward function. The DRL framework in the lower level learns the locomotion control policy using the learned reward function. Three sets of experiments are conducted to verify the effectiveness of the proposed approach: training soft actor-critic and proximal policy optimization agents in MuJoCo environments, training the proximal policy optimization agent in a humanoid robot environment built with Isaac Lab, and transferring the agent to the real-world Unitree G1 humanoid robot through sim-to-real. The experimental results demonstrate that the proposed automatic reward learning method substantially improves learning efficiency and achieves superior performance to manually designed reward functions in both simulation and real-world deployment. By enhancing the success rate of transferring control policies from simulation to real-world humanoid robots, this approach provides a promising pathway toward accelerating the deployment of stable and adaptive humanoid robots in practical applications.

Identifiers

PMID42495013
PMCPMC13395157

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.