Master's Thesis on Human–Robot Interaction Design
July 2021
Overview
For large human-operated robots, surrounding sensors alone cannot distinguish an intentional approach to a work target from an unintended approach to an overlooked obstacle. I designed a collision-prevention system that treats the operator’s field of view as a signal of “permitted to touch”, slowing or stopping the arm as it approaches objects outside that field.
Objects do not need to be registered in advance as “operable” or “protected”; the distinction changes with the operator’s gaze during the task. I implemented the system on a 1.9-metre-tall humanoid heavy machinery robot and confirmed the intended behaviour across five demonstration trials. The system in action can be seen in the demo video.
As a master's thesis in Design & Engineering at Politecnico di Milano, it was undertaken during a research internship at Jinki Ittai Co., Ltd., and presented at The 39th Annual Conference of the Robotics Society of Japan (RSJ2021).
Problem — Safety configuration can constrain field work
Human-operated robots are needed in places that are not equipped for autonomous robots. In locations like construction sites or disaster areas, the arrangement of people and equipment changes, and what can be touched also varies with each task. To use the great power of machines, it is necessary to be able to touch the required objects while preventing unintended contact.
Existing collision avoidance methods include covering the robot with proximity sensors or reproducing the surroundings in a virtual space using depth sensors. However, in both cases, it is necessary to pre-classify objects as either 'operable' or 'protected.' If nothing is registered, the robot cannot touch anything for safety, and registering everything requires setup and calibration for each site. This preparation burden has reduced the practicality of robots used in multiple locations.
Idea — Turn gaze into a condition for touch
In the control loop of a human-operated robot, there is already a person who judges the situation. I thought that if we could pass to the machine where the person is focusing their attention, it might be possible to estimate the intention of contact during the task instead of pre-registering the targets.
Therefore, the operator's central field of vision is used as the boundary. Objects within the field of vision are considered 'operable objects' that the operator is aware of. Objects outside of it are considered 'protected objects' that might be overlooked. Regardless of the position of their gaze, people are always considered protected objects.
Gaze is not used to issue commands. Instead, the system translates the operator’s cognitive state into conditions that determine where the robot is allowed to move.
Interaction design
The system does not treat the gaze position as a simple on-and-off switch. When approaching an object that is not being looked at, the arm gradually slows down according to the distance and stops before making contact. After stopping, if the operator directs their gaze toward the object, the restrictions are gradually lifted. This avoids sudden stops and sudden movements, allowing the operator to anticipate changes.
On the other hand, even if you temporarily look away from the object you are already touching, ongoing work is not stopped. This is because if grasping or operation is interrupted just by averting your gaze, it actually becomes unstable. Only approaching people is treated separately, and it is restricted regardless of whether the operator is watching or not.
The aim was not to uniformly slow down operations for safety. It is to maintain the robot's capabilities within the operator's grasp and to protect only the areas that are beyond their attention, thereby achieving both safety and operability.
Intended use contexts
The focus was on situations where large forces are handled close to people, such as construction sites or disaster sites. In such places, it is difficult to completely separate the surroundings with fences or to pre-register all objects. A system that requires redoing safety settings every time the arrangement changes would lose both urgency and versatility.
Distinctions based on gaze use a person's situational awareness at the moment instead of a fixed environmental model. This may reduce the burden of site preparation while limiting approaches to overlooked targets. However, what was confirmed in this study was only that the system operates as defined, and the effect on accident reduction has not been evaluated in practice.
Prototype and preliminary evaluation
Implementation
The system was implemented in a robot with the shape of a human upper body, 1.9 meters tall, weighing 600 kg, and having 30 degrees of freedom. The operator wears a VR headset and controls it by synchronizing their movements with the robot's field of view and head orientation.
The workspace was captured as a point cloud using a LIDAR on the robot and represented in the virtual space as voxels with collision detection. A digital twin that synchronizes this space with the real machine operates. The operator's central field of view was represented as a 60°×55° cone, and the orientation was approximated using the front direction of the headset. When the arm approaches a voxel outside the field of view, a warning is sent to the real machine, causing it to slow down or stop.
Preliminary evaluation
The trials followed a sequence in which the operator approached a target while looking away, looked back at it to release the restriction after the arm stopped, and then grasped the ball. Four of the five trials used an experienced operator; the fifth used someone in the cockpit for only the second time. In every trial, the arm stopped when the operator looked away and resumed when they looked at the target. This was not a user study of effectiveness, but a preliminary evaluation of whether the designed behaviour worked on the physical robot.
Limitations and next steps
The prototype does not implement human detection. A rule was set to always protect people regardless of gaze, but in a real environment, a mechanism to reliably identify people is necessary. Also, because the front direction of the headset was used as a proxy value rather than the gaze itself, movements of just the eyes cannot be detected.
If you use a stationary camera or smart glasses with gaze tracking, the conditions for the wearable device and estimation accuracy will change. When deploying to a mobile robot, point cloud processing that can follow the robot's own position changes will also be necessary.
The evaluation is limited to qualitative verification through five demonstration trials. To assess safety, work speed, and operator burden, a quantitative comparison with conventional methods under actual working conditions is necessary.
What this research changed in my design practice
In this study, we treated a person's gaze not as an input device, but as a condition for allowing machine actions. Depending on what a person is currently aware of, the robot's behavior toward the same object changes. What we designed was not individual actions, but the relationship between a person's cognitive state and the machine's authority.
This way of thinking continues in my design philosophy: understand a person’s state, translate it into constraints the system must uphold, and let machines or AI exercise their capabilities within those boundaries. This master’s research became the starting point for that approach.
Acknowledgments
This research was conducted during an internship at Jinki Ittai Co., Ltd. as part of the Master's program in Design & Engineering at Politecnico di Milano. I am deeply grateful to both organizations for making this opportunity possible.
Takafumi Horiuchi, Katsuya Kanaoka, “Gaze-based Unintended Contact Prevention for Human-Operated Robots,” The 39th Annual Conference of the Robotics Society of Japan, 2021, 2J2-07. Conference programme














