On 19 June 2026, a joint research team from the Collaborative and Versatile Robots Laboratory (led by Prof. Fei Chen) in CUHK’s Department of Mechanical and Automation Engineering and the Hong Kong Centre for Logistics Robotics received the Best Paper Award at the 20th IEEE International Conference on Control and Automation (ICCA 2026) in Almaty, Kazakhstan.
Paper Title: Towards Efficient Robot Learning: Diffusion-style Skill Learning and Transfer on Platform with Multi-modal Perception and Force Feedback
Authors:
Dianxi Li (Ph.D. student, Department of Mechanical and Automation Engineering, CUHK)
Zhipeng Dong (Co-founder and Chief Operating Officer, SOTA Robotics (Hong Kong) Limited)
Zhuo Li (Ph.D. student, Department of Mechanical and Automation Engineering, CUHK)
Wenrui Liu (Ph.D. student, Department of Mechanical and Automation Engineering, CUHK)
Fei Chen (Assistant Professor, Department of Mechanical and Automation Engineering, CUHK)
Project Description:
From High-quality Demonstration to Real-world Execution
Bimanual robots operating in cooking, assembly, cleaning, and service tasks must coordinate two arms while managing physical contact, precise poses, and multi-stage actions. Motion trajectories alone are not sufficient for teaching these tasks. Demonstrators also need to understand how robots interact with the environment. Without force feedback, contact state must be inferred largely from vision, making it harder to adjust motion and force naturally during demonstration.
Real deployments also require more than repeating an isolated action. Robots must learn reusable skills from multimodal demonstrations, adapt prior knowledge to new tasks, preserve smooth action continuity, and decide when to switch between skills. When users express goals in natural language, the system must also interpret intent, plan over a long horizon, and schedule the appropriate low-level policies. The award-winning work addresses these connected needs through a complete pipeline spanning force-aware teaching, skill learning, intelligent scheduling, and real-world deployment.
Highlight 1: Force-feedback Bimanual Teaching and Multimodal Data
The team built a leader-follower platform using four Franka arms. Two leader arms are operated by the human demonstrator, while two follower arms execute the task in the workspace. Joint-torque control on both sides enables real-time force feedback, allowing the demonstrator to feel contact onset, resistance, and manipulation stability instead of relying on visual feedback alone.
The follower side records global and wrist-view images, depth, three-dimensional point clouds, six-axis end-effector force/torque, joint states, end-effector poses, and gripper states. Foot pedals support gripper operation, recording control, keyframe annotation, and task-completeness labeling. The resulting demonstrations retain synchronized visual, geometric, action, and physical-interaction information.
Highlight 2: Efficient Diffusion-style Skill Learning and Transfer
A multi-stage point-cloud encoder compresses raw three-dimensional observations into a compact global representation. This representation is combined with vision, force, and robot proprioception to condition the diffusion-style policy, providing spatial information for bimanual action generation without burdening the policy with the full raw point cloud.
For skill transfer, target policies do not always start from a standard Gaussian prior. Instead, an already learned source policy provides a starting distribution closer to the target behavior, and a stochastic interpolation bridge refines samples toward the target skill. This mechanism allows prior skill knowledge to support the expansion and refinement of the robot skill library.
The work also addresses discontinuities between successive action chunks. Previously executed actions were fixed as known conditions, while only the unknown future horizon is generated through action-sequence inpainting. A task-completeness signal is predicted alongside the action sequence; once it crosses a threshold, the system can advance to the next skill. Together, these mechanisms support smoother motion and more reliable long-horizon composition.

Leader arm, network service, and Curi follower modules
Highlight 3: LLM-guided Skill-library Execution
The platform connects speech interaction, a large language model, and a library of learned robot skills. Spoken requests are transcribed into text. Separate language-model contexts generate a natural response and an executable task plan. The planner interprets the user’s intent, considers the current scene state, decomposes the goal into subtasks, and calls the corresponding policy interfaces. After each skill is completed, the system reassesses the state and selects the next action until the overall task is finished.
Validation in a Congee-shop Scenario
The team evaluated the system in a realistic congee-shop setting with skills for dispensing ingredients, returning the ingredient cup, handling the pot lid, adding water, cooking and serving congee, and wiping the table. These tasks combine bimanual coordination, tool use, narrow-space positioning, waiting phases, liquid handling, and physical contact, providing a representative long-horizon test environment.
When a user requests congee in natural language, the planner determines the required sequence from the current scene state and invokes the appropriate skills. During serving, the two arms coordinate the pot lid, ladle, and pot. When bowl positions or quantities change, the learned policy adjusts its motion to the new state, while the action-continuity mechanism keeps the ladle trajectory stable and reduces abrupt changes during liquid handling.
Research Significance
Beyond a single congee-shop demonstration, the work integrates high-quality force-aware teaching, multimodal representation, transferable skill learning, continuous action generation, task-completeness estimation, and language-guided scheduling in one system. It offers a practical path for learning complex bimanual skills more efficiently and composing them into deployable long-horizon tasks.
The IEEE ICCA 2026 Best Paper Award recognizes the joint team’s work in robot learning, bimanual manipulation, and real-world deployment. The CUHK Collaborative and Versatile Robots Laboratory and the Hong Kong Centre for Logistics Robotics will continue to advance artificial intelligence and robotics research and support the translation of these technologies into service, manufacturing, and other real environments.
