MAMHOI: Factorizing Scene-Aware Human–Object Interaction Through Affordances

An affordance-mediated formulation that connects scene-level interaction feasibility with detailed human–object motion synthesis.

Mingyuan Lei Yoonchang Sung Tat-Jen Cham

College of Computing and Data Science
Nanyang Technological University, Singapore

Abstract

Generating realistic human–object interactions in complex 3D scenes requires two complementary capabilities: reasoning about interaction feasibility in the environment and synthesizing realistic human–object motion. Supervision for these capabilities, however, is rarely available jointly at scale.

We present MAMHOI, an affordance-mediated factorization for scene-aware human–object interaction generation. A scene-conditioned model first predicts where and how an interaction can be feasibly executed. An affordance-conditioned HOI model then generates the corresponding synchronized human–object motion. This formulation enables scene understanding and interaction dynamics to be learned from complementary sources of supervision without requiring paired human–object–scene data.

Experiments in complex indoor environments show that MAMHOI reduces object–scene penetration while preserving human–object interaction quality, producing more realistic and physically feasible scene-aware interactions.

MAMHOI transfers scene-free human-object interaction priors to complex scenes through scene-aware motion affordances

MAMHOI transfers scene-free HOI priors to complex environments through scene-aware motion affordances.

Video

The video presents the formulation, model behavior, and representative results in complex indoor scenes.

Qualitative Comparison

Generated interactions under identical scene, object, and task conditions.

Motivation

Scene-aware HOI generation must determine both where an interaction is physically feasible and how the human and object should move together. Yet the strongest supervision for these two capabilities comes from different datasets.

A

Scene understanding

Human–scene datasets provide environment-aware motion but lack detailed object manipulation dynamics.

B

Interaction synthesis

Human–object datasets capture rich contact and manipulation behavior but are typically scene-free.

C

Modeling objective

A useful representation must transfer scene constraints without overwriting the interaction distribution learned from motion-rich HOI data.

Method

MAMHOI separates scene understanding from human–object motion synthesis through a sequence-level motion affordance.

MAMHOI framework with scene-to-affordance and affordance-to-HOI-motion factors
FACTOR I · SCENE → AFFORDANCE

Scene understanding

An affordance denoising network predicts task-specific feasible regions from scene geometry, the planned object path, and the language instruction.

  • Learned from indoor HSI data
  • Task-relevant spatial constraints
  • Sequence-level motion affordance
FACTOR II · AFFORDANCE → HOI MOTION

Affordance-guided synthesis

Local occupancy around the human, object, and goal is encoded as affordance features that condition a diffusion-based human–object motion generator.

  • Learned from scene-free HOI data
  • Contact-aware interaction dynamics
  • Segment-wise scene adaptation

Citation

If you find this work useful, please cite:

@misc{lei2026mamhoifactorizingsceneawarehumanobject,
  title={MAMHOI: Factorizing Scene-Aware Human-Object Interaction through Affordances},
  author={Mingyuan Lei and Yoonchang Sung and Tat-Jen Cham},
  year={2026},
  eprint={2610.12416},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2610.12416}
}