TY - GEN
T1 - Discounted Sampling Policy Gradient for Robot Multi-objective Visual Control
AU - Xu, Meng
AU - Zhang, Qingfu
AU - Wang, Jianping
PY - 2021/3/28
Y1 - 2021/3/28
N2 - Robot visual control often involves multiple objectives such as achieving high efficiency, maintaining stability, and avoiding failure. This paper proposes a novel Vision-Based Control method (VBC) with the Discounted Sampling Policy Gradient (DSPG) and Cosine Annealing (CA) to achieve excellent multi-objective control performance. In our proposed visual control framework, a DSPG learning agent is employed to learn a policy estimating continuous kinematics for VBC. The deep policy maps the visual observation to a specific action in an end-to-end manner. The DSPG agent finally can update the policy to obtain the optimal or near-optimal solution using shaped rewards from the environment. The proposed VBC-DSPG model is optimized using a heuristic method. Experimental results demonstrate that the proposed method performs very well compared with some classical competitors in the multi-objective visual control scenario.
AB - Robot visual control often involves multiple objectives such as achieving high efficiency, maintaining stability, and avoiding failure. This paper proposes a novel Vision-Based Control method (VBC) with the Discounted Sampling Policy Gradient (DSPG) and Cosine Annealing (CA) to achieve excellent multi-objective control performance. In our proposed visual control framework, a DSPG learning agent is employed to learn a policy estimating continuous kinematics for VBC. The deep policy maps the visual observation to a specific action in an end-to-end manner. The DSPG agent finally can update the policy to obtain the optimal or near-optimal solution using shaped rewards from the environment. The proposed VBC-DSPG model is optimized using a heuristic method. Experimental results demonstrate that the proposed method performs very well compared with some classical competitors in the multi-objective visual control scenario.
KW - Multi-objective visual control
KW - Kinematics
KW - Discounted sampling policy gradient
KW - Cosine annealing
UR - https://www.scopus.com/pages/publications/85107285289
UR - https://www.scopus.com/record/pubmetrics.uri?eid=2-s2.0-85107285289&origin=recordpage
U2 - 10.1007/978-3-030-72062-9_35
DO - 10.1007/978-3-030-72062-9_35
M3 - RGC 32 - Refereed conference paper (with host publication)
SN - 9783030720612
T3 - Lecture Notes in Computer Science
SP - 441
EP - 452
BT - Evolutionary Multi-Criterion Optimization
A2 - Ishibuchi, Hisao
A2 - Zhang, Qingfu
A2 - Cheng, Ran
A2 - Li, Ke
A2 - Li, Hui
A2 - Wang, Handing
A2 - Zhou, Aimin
PB - Springer
CY - Cham
T2 - 11th International Conference on Evolutionary Multi-Criterion Optimization (EMO 2021)
Y2 - 28 March 2021 through 31 March 2021
ER -