ArticleResearch (Washington, D.C.)2026
Deep Reinforcement Learning for Real-World Humanoid Robot Locomotion Control with Automatic Reward Learning.
Article in Research (Washington, D.C.), 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
10 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Humanoid robots possess the potential to solve complex problems across diverse environments, such as nuclear-contaminated zones, epidemic-affected areas, and extraterrestrial missions. However, a humanoid robot is an inherently complex system that integrates multiple disciplines, including perception, mechanical design, materials science, and motion control, each of which requires comprehensive and in-depth investigation. Among these aspects, motion control plays a crucial role, as it directly determines the robot's motion accuracy, stability, and flexibility. In recent years, with the rapid evolution of graphics-processing-unit-based parallel computing and high-fidelity simulation environments, various deep-reinforcement-learning (DRL)-based approaches have been proposed to achieve precise and robust motion control due to its flexibility and adaptability in uncertain and dynamic environments. However, the inherent complexity and uncertainty of real-world tasks pose substantial challenges when designing effective reward functions for DRL agents. Most current methods typically rely on manually engineered or externally tuned reward signals and therefore require considerable domain expertise, associated with considerable human efforts and a long convergence time; these issues may even trigger mission failure. This work proposes an automatic reward learning method to derive reward functions for DRL in humanoid robot locomotion control. Specifically, a bilevel optimization framework is developed to enable automatic reward learning during policy learning. The reward learning mechanism in the upper level adaptively constructs and optimizes the reward function. The DRL framework in the lower level learns the locomotion control policy using the learned reward function. Three sets of experiments are conducted to verify the effectiveness of the proposed approach: training soft actor-critic and proximal policy optimization agents in MuJoCo environments, training the proximal policy optimization agent in a humanoid robot environment built with Isaac Lab, and transferring the agent to the real-world Unitree G1 humanoid robot through sim-to-real. The experimental results demonstrate that the proposed automatic reward learning method substantially improves learning efficiency and achieves superior performance to manually designed reward functions in both simulation and real-world deployment. By enhancing the success rate of transferring control policies from simulation to real-world humanoid robots, this approach provides a promising pathway toward accelerating the deployment of stable and adaptive humanoid robots in practical applications.
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.