Takafumi Horiuchi

Master's Thesis on Human–Robot Interaction Design

Overview

For large human-operated robots, surrounding sensors alone cannot distinguish an intentional approach to a work target from an unintended approach to an overlooked obstacle. I designed a collision-prevention system that treats the operator’s field of view as a signal of “permitted to touch”, slowing or stopping the arm as it approaches objects outside that field.

Objects do not need to be registered in advance as “operable” or “protected”; the distinction changes with the operator’s gaze during the task. I implemented the system on a 1.9-metre-tall humanoid heavy machinery robot and confirmed the intended behaviour across five demonstration trials. The system in action can be seen in the demo video.

As a master's thesis in Design & Engineering at Politecnico di Milano, it was undertaken during a research internship at Jinki Ittai Co., Ltd., and presented at The 39th Annual Conference of the Robotics Society of Japan (RSJ2021).


  1. Title slide. Awareness-Based Collision Prevention for Human-Controlled Robotic Manipulation, Politecnico di Milano Design & Engineering Master's Program, 2020–2021.
  2. Background. The thesis topic was decided during a research internship at Jinki Ittai Co., Ltd., a company that develops humanoid heavy machinery and operates it using a master-slave system.
  3. Comparison between autonomous robots and human-operated robots. The former work efficiently in designed spaces inside fences, while the latter work with humans in uncertain spaces.
  4. Statistics from Europe, the United States, and Japan indicate that contact with moving objects accounts for a significant proportion of construction site accidents.
  5. Previous research on collision avoidance. Tactile skin covering the manipulator and a depth sensor that virtually reproduces the surroundings.
  6. Current issue. The existing system requires predefining the target as 'protected' or 'operable' using attached tags, image analysis, and a trained classifier.
  7. Proposal. Conventional systems only use passive sensor information, but in human-operated robots, there is inherently a person within the control loop, and their gaze can be used as an active parameter.
  8. Classification rules. Objects within the operator's central field of vision are operable, while objects outside are protected.
  9. An overall view of the prototype. It combines a real robot with a virtual computational space, followed by three technological elements.
  10. Element ①. Represent the central visual field as a pyramid with a 60°×55° cone and approximate its orientation with the front direction of the VR headset.
  11. Element ②. The LIDAR mounted on the robot's head scans the surroundings as a point cloud, with each point represented as a voxel that has a collider.
  12. Element 3. The robot's digital twin moves in the same way as the actual machine, and when its arm touches a voxel outside the field of view, a proximity warning is issued.
  13. Proof-of-concept demo. A still image of a robot reaching for a ball on a pole, and a voxel display in virtual space.
  14. Future challenges. Detecting people within the workspace, verifying other methods for measuring gaze direction, and conducting quantitative evaluations in the field.
  15. Summary of the conclusion. The research theme, the value of the methodology, novelty, and how the system was verified.

Problem — Safety configuration can constrain field work

Human-operated robots are needed in places that are not equipped for autonomous robots. In locations like construction sites or disaster areas, the arrangement of people and equipment changes, and what can be touched also varies with each task. To use the great power of machines, it is necessary to be able to touch the required objects while preventing unintended contact.

Existing collision avoidance methods include covering the robot with proximity sensors or reproducing the surroundings in a virtual space using depth sensors. However, in both cases, it is necessary to pre-classify objects as either 'operable' or 'protected.' If nothing is registered, the robot cannot touch anything for safety, and registering everything requires setup and calibration for each site. This preparation burden has reduced the practicality of robots used in multiple locations.

Idea — Turn gaze into a condition for touch

In the control loop of a human-operated robot, there is already a person who judges the situation. I thought that if we could pass to the machine where the person is focusing their attention, it might be possible to estimate the intention of contact during the task instead of pre-registering the targets.

Therefore, the operator's central field of vision is used as the boundary. Objects within the field of vision are considered 'operable objects' that the operator is aware of. Objects outside of it are considered 'protected objects' that might be overlooked. Regardless of the position of their gaze, people are always considered protected objects.

Gaze is not used to issue commands. Instead, the system translates the operator’s cognitive state into conditions that determine where the robot is allowed to move.

View from above. In front of the robot, the operator's central vision is depicted as a cone. Target A, which is inside the cone, is shown as a white circle, while targets B, C, and D, which are outside, are represented by dark squares. Two people are standing inside and outside the cone.
Target A, which is within the operator's central field of vision, can be operated, while targets B, C, and D, which are outside, are to be protected. People are protected regardless of where they are standing.

Interaction design

The system does not treat the gaze position as a simple on-and-off switch. When approaching an object that is not being looked at, the arm gradually slows down according to the distance and stops before making contact. After stopping, if the operator directs their gaze toward the object, the restrictions are gradually lifted. This avoids sudden stops and sudden movements, allowing the operator to anticipate changes.

On the other hand, even if you temporarily look away from the object you are already touching, ongoing work is not stopped. This is because if grasping or operation is interrupted just by averting your gaze, it actually becomes unstable. Only approaching people is treated separately, and it is restricted regardless of whether the operator is watching or not.

The aim was not to uniformly slow down operations for safety. It is to maintain the robot's capabilities within the operator's grasp and to protect only the areas that are beyond their attention, thereby achieving both safety and operability.

A flowchart that starts from a state of free movement. It branches based on whether an object is nearby, whether it is a person, whether the manipulator is in the direction of approaching the person, or whether the object is in the gaze area, leading to precise operations, gradual stopping, or locking.
In the decision-making process, the node that reads the operator's active parameters is only the one for 'whether the target is in the gaze area.' All other branches operate on passive sensor information.

Intended use contexts

The focus was on situations where large forces are handled close to people, such as construction sites or disaster sites. In such places, it is difficult to completely separate the surroundings with fences or to pre-register all objects. A system that requires redoing safety settings every time the arrangement changes would lose both urgency and versatility.

Distinctions based on gaze use a person's situational awareness at the moment instead of a fixed environmental model. This may reduce the burden of site preparation while limiting approaches to overlooked targets. However, what was confirmed in this study was only that the system operates as defined, and the effect on accident reduction has not been evaluated in practice.

Prototype and preliminary evaluation

Implementation

The system was implemented in a robot with the shape of a human upper body, 1.9 meters tall, weighing 600 kg, and having 30 degrees of freedom. The operator wears a VR headset and controls it by synchronizing their movements with the robot's field of view and head orientation.

The workspace was captured as a point cloud using a LIDAR on the robot and represented in the virtual space as voxels with collision detection. A digital twin that synchronizes this space with the real machine operates. The operator's central field of view was represented as a 60°×55° cone, and the orientation was approximated using the front direction of the headset. When the arm approaches a voxel outside the field of view, a warning is sent to the real machine, causing it to slow down or stop.

A scene in a virtual space. Yellow voxels reconstruct the surroundings, and a translucent red area labeled 'operator's central field of view' pierces through it, while a white digital twin robot reaches out to an obstacle outside that area.
A virtual space containing a robot and the scanned point cloud. Here, the robot is touching obstacles that are outside the operator's central field of vision.

Preliminary evaluation

The trials followed a sequence in which the operator approached a target while looking away, looked back at it to release the restriction after the arm stopped, and then grasped the ball. Four of the five trials used an experienced operator; the fifth used someone in the cockpit for only the second time. In every trial, the arm stopped when the operator looked away and resumed when they looked at the target. This was not a user study of effectiveness, but a preliminary evaluation of whether the designed behaviour worked on the physical robot.

A photo of a robot. The operator is riding in the upper frame and is facing sideways. The robot hand stops in front of a blue and yellow ball on a yellow pole, and the remaining distance is indicated.
The operator's gaze is diverted, and the arm stops at a distance from the target.
A photo from a low angle. The operator is looking at the ball, and the robot hand is grabbing the blue and yellow ball on top of the yellow pole.
If your gaze is directed at the target, the same movement can proceed without restriction.

Limitations and next steps

The prototype does not implement human detection. A rule was set to always protect people regardless of gaze, but in a real environment, a mechanism to reliably identify people is necessary. Also, because the front direction of the headset was used as a proxy value rather than the gaze itself, movements of just the eyes cannot be detected.

If you use a stationary camera or smart glasses with gaze tracking, the conditions for the wearable device and estimation accuracy will change. When deploying to a mobile robot, point cloud processing that can follow the robot's own position changes will also be necessary.

The evaluation is limited to qualitative verification through five demonstration trials. To assess safety, work speed, and operator burden, a quantitative comparison with conventional methods under actual working conditions is necessary.

What this research changed in my design practice

In this study, we treated a person's gaze not as an input device, but as a condition for allowing machine actions. Depending on what a person is currently aware of, the robot's behavior toward the same object changes. What we designed was not individual actions, but the relationship between a person's cognitive state and the machine's authority.

This way of thinking continues in my design philosophy: understand a person’s state, translate it into constraints the system must uphold, and let machines or AI exercise their capabilities within those boundaries. This master’s research became the starting point for that approach.

Acknowledgments

This research was conducted during an internship at Jinki Ittai Co., Ltd. as part of the Master's program in Design & Engineering at Politecnico di Milano. I am deeply grateful to both organizations for making this opportunity possible.

Takafumi Horiuchi, Katsuya Kanaoka, “Gaze-based Unintended Contact Prevention for Human-Operated Robots,” The 39th Annual Conference of the Robotics Society of Japan, 2021, 2J2-07. Conference programme