Synthetically Trained Neural Networks for Learning Human-Readable Plans from Real-World Demonstrations

We present a system to infer and execute a human-readable program from a real-world demonstration. The system consists of a series of neural networks to perform perception, program generation, and program execution. Leveraging convolutional pose machines, the perception network reliably detects the bounding cuboids of objects in real images even when severely occluded, after training only on synthetic images using domain randomization. To increase the applicability of the perception network to new scenarios, the network is formulated to predict in image space rather than in world space. Additional networks detect relationships between objects, generate plans, and determine actions to reproduce a real-world demonstration. The networks are trained entirely in simulation, and the system is tested in the real world on the pick-and-place problem of stacking colored cubes using a Baxter robot.

Authors

Jonathan Tremblay

Thang To (NVIDIA)

Artem Molchanov (NVIDIA, USC)

Stephen Tyree

Jan Kautz

Stan Birchfield

Publication Date

Monday, May 21, 2018

Published in

IEEE International Conference on Robotics and Automation (ICRA) 2018

Research Area

Computer Vision

Robotics

External Links

arXiv paper

video

NVIDIA developer news blog

Uploaded Files

Paper8.48 MB

Copyright

This material is posted here with permission of the IEEE. Internal or personal use of this material is permitted. However, permission to reprint/republish this material for advertising or promotional purposes or for creating new collective works for resale or redistribution must be obtained from the IEEE by writing to pubs-permissions@ieee.org.