A dataset & benchmark for human–robot interaction

Robots see the room.
Humans give it meaning.

HEIR — Harness Egocentric Intent for Human–Robot Interactions

Connecting what people see, look at, and say with how robots navigate, manipulate, and collaborate in everyday spaces.

Dataset on Hugging Face Explore the demos
38.56h

Recorded interaction

300

Continuous sessions

1,836

Request-level episodes

6,416

Execution subtasks

17

Participants

HUMAN CONTEXT ROBOT EXECUTIONONE SHARED INTERACTION
Three consecutive requests in one HEIR session: fetch salad dressing, hold it during use, and return it. Human view and gaze are paired with navigation, manipulation and spatial trajectories.

01 / The missing perspective

“Bring that over.”
What does that mean?

A request is more than a sentence. Its meaning depends on a person's attention, activity, and the interactions that came before it.

HEIR preserves both sides of the interaction: human egocentric video, measured gaze, and speech alongside robot observations and teleoperated navigation–manipulation trajectories.

Egocentric view Measured gaze Speech

02 / In everyday spaces

From a glance
to a shared task.

Two recorded interactions bring human intent and mobile robot execution into the same view.

01 · OBJECT RETRIEVAL

“Please bring that object over for me.”

Resolve the target. Bring it closer.

A situated request leads into object retrieval, movement across the room, and handover. The multi-view recording shows the interaction from both human and robot perspectives.

Referential intentNavigation + manipulation

Edited clip · 1:37 · Speed changes are marked in the video.

02 · COLLABORATIVE MANIPULATION

“Let's smooth out the bedsheet together.”

One task. Two participants.

A person and a dual-arm robot grasp and adjust the same bedsheet, illustrating the physical context surrounding a collaborative household request.

Human–robot collaborationDeformable object

Edited clip · 0:44 · Speed changes are marked in the video.

Qualitative interaction examples. These clips illustrate the tasks; quantitative success rates are reported separately below.

03 / The HEIR dataset

Two perspectives.
A continuous story.

Requests unfold within sessions, so later episodes can refer to objects and states established earlier.

DATA HIERARCHY
S

Session

A continuous household interaction

E

Episode

One natural human request

T

Subtask

Navigation or manipulation

HUMAN SIDE

Intent in context

  • Egocentric scene video
  • Measured eye gaze
  • Spoken requests and interaction cues

Pupil Labs Neon eye-tracking glasses

ROBOT SIDE

Perception & action

  • Two top and two wrist camera streams
  • Arm, gripper, torso, and base states
  • Teleoperated navigation–manipulation trajectories

Galaxea R1 Lite mobile dual-arm platform

ANNOTATIONS

From request to execution

  • Grounded intent, objects, and destinations
  • Nominal plans and realized subtask boundaries
  • Subgoals and completion conditions

Human annotations with VLM-assisted refinement

HEIR collection overview: human and robot sensing, staged multi-episode sessions, and temporally grounded human annotations followed by VLM-assisted refinement.
Aligned multimodal recordings link human intent, task semantics, and robot behavior.
240 / 60

Training / testing sessions
Split at the session level

7.71 min

Average session duration

6.12

Episodes per session, on average

3.49

Subtasks per episode, on average

04 / A hierarchical benchmark

Understanding is only
the beginning.

Evaluate each component, then examine what happens when they work together in the real world.

WHAT DOES THE HUMAN MEAN?

Ground the request.

Use speech, measured gaze, human and robot views, and interaction history to infer a grounded task description, target objects, destinations, and a nominal execution plan.

Metrics: Slot F1 · Location accuracy · Plan exact match · Entity–visibility F1

STRUCTURED INTENT · ILLUSTRATIVE

Request“Please bring that object over for me.”
ResolveObject · destination · interaction context
PlanNavigate → retrieve → navigate → hand over
Human contextIntent reasonerClosed-loop plannerVLN / VLA execution

The planner chooses Continue, Switch, or Done from current observations and execution history.

05 / What the benchmark reveals

Human context helps.
Long horizons remain hard.

Results from the manuscript. Offline component scores and real-robot success measure different parts of the problem.

INTENT GROUNDING+17.42pp

Richer context, clearer intent.

Slot F1 rises from 60.38 with text only to 77.80 with human and robot views plus measured gaze.

Gemma-4-31B · no interaction history · Table 2
NAVIGATION ADAPTATION90.69%

Learning the household setting.

Fine-tuned InternVLA-N1 reaches 90.69% offline action accuracy, compared with 49.08% for its original checkpoint.

Offline navigation action prediction · Table 5
END-TO-END EXECUTION18.52%

Good components are not enough.

The integrated system succeeds on 5 of 27 episodes, revealing error accumulation across long-horizon interactions.

Real-robot episode success · Section 5.5

EXPLORE THE RESULTS

Human context resolves ambiguity.

Gemma-4-31B input ablations without interaction history. Scores are percentages.
Input conditionSlot F1 ↑Location ↑Plan EM ↑RV F1 ↑BERTScore ↑
Text only60.3819.4233.24N/A92.96
Robot view71.3955.3864.3261.6393.19
Human + robot75.7472.1873.5160.9493.27
Human + robot + center prior75.2470.0878.9264.1093.23
Human + robot + gaze77.8077.9579.1963.9593.23

Table 2. Measured gaze improves grounding beyond a center prior; the center-prior condition is slightly higher on entity–visibility F1.

THE REAL-ROBOT GAP

From individual subtasks
to a complete request.

End-to-end success requires intent, planning, navigation, and manipulation to stay aligned throughout the episode.

Navigation subtasks InternVLA-N1 · 26 / 3281.25%
Manipulation subtasks Lingbot-VLA2 · 28 / 3971.79%
Complete episodes Integrated system · 5 / 2718.52%

06 / Research resources

Build toward robots
that understand people.

Explore the benchmark data and task definitions behind HEIR. The paper and code are coming soon.

The manuscript describes planned releases of annotations, processed evaluation instances, prompts, and evaluation code. Sensitive raw recordings are intended for controlled access after privacy review.

Click outside the figure or press Escape to close.