How to Learn from What a Human Would Avoid?

Intervention-Aware World Models with Real-World RL for Dexterous Manipulation

Jiaju Yin1,*, Zhenhui Zhang1,*, Lixin Xu1, Heng Zhang2, Jun Shao3, Yating Feng1, Arash Ajoudani2,†, Renjing Xu1,†

1HKUST (Guangzhou) 2Italian Institute of Technology 3Zhejiang University

*Equal contribution Corresponding author

WHIRL turns human takeovers into predictive intervention signals for real-world residual RL with a 16-DoF dexterous hand.

WHIRL teaser showing intervention-aware world model and real robot tasks

Method

Intervention-aware residual RL

WHIRL treats each human takeover as supervision about states the operator would avoid, then uses that signal to train a residual policy more proactively.

WHIRL main pipeline from teleoperation and interventions to residual reinforcement learning

Main pipeline

A behavior prior is first learned from teleoperated demonstrations. During real-world RL, the residual policy acts by default, while operator takeovers provide binary intervention labels that are reused as training signal instead of being discarded as only online corrections.

Intervention-aware latent world model with dynamics, reward, termination, and intervention heads

Intervention-aware world model

The latent world model keeps standard dynamics, reward, and termination heads for critic learning, and adds an intervention-probability head. That prediction becomes an actor-side risk-shaping term, steering the policy away from intervention-prone states before the human needs to take over.

Real-world rollouts

Five dexterous manipulation tasks

Short-horizon tasks

Long-horizon task

Long Horizon

Real-world rollout for the long-horizon drawer-place-close sequence.

Conclusion

Human takeovers become reusable risk predictions

WHIRL uses pedal interventions as forward supervision for an intervention-aware world model. The standard dynamics, reward, and termination heads support the critic, while the intervention head shapes the actor away from states that are likely to need human control. Across five real-world dexterous manipulation tasks, this reduces operator-controlled training steps while improving autonomous success.

Citation

@article{whirl2026,
  title   = {How to Learn from What a Human Would Avoid? Intervention-Aware World Models with Real-World RL for Dexterous Manipulation},
  author  = {Yin, Jiaju and Zhang, Zhenhui and Xu, Lixin and Zhang, Heng and Shao, Jun and Feng, Yating and Ajoudani, Arash and Xu, Renjing},
  journal = {arXiv preprint},
  year    = {2026}
}