ORGA
ORGA Architecture
Object-centric extraction of transferable geometric constraints and spatial structure for zero-shot manipulation
The real world cannot be exhausted by data
Long-tail scenes
Pose, color, shape, and lighting vary endlessly — exhaustive collection never catches up
High deployment cost
Every new scene demands fresh data and retraining, stretching timelines
Trajectory memorization
Classic imitation learning remembers seen motions and fails on novel objects
ORGA: Object-centric Representation for Generalization
Extract latent geometric constraints and spatial structure from few demonstrations — so robots move from memorizing trajectories to understanding transferable manipulation principles
01
Across objects
Learn to peel a cucumber — then peel a carrot
02
Across scenes
New poses, lighting, and layouts — no retraining
03
Across embodiments
One model teaches the same skill to different robots
01
Object-centric representation
Anchor manipulation understanding in object geometry and interaction — not absolute trajectories in camera space
02
Geometric constraint extraction
Distill latent spatial constraints and structural priors from few demonstrations into reusable manipulation knowledge
03
Zero-shot transfer
Generalize to novel objects, scenes, and embodiments without per-instance retraining
From human-centric multimodal capture to cross-scene expansion and cross-embodiment retargeting — how one skill understanding transfers to new environments and robot platforms
Data capture
Human-centric multimodal teaching
Wearable sensing records vision, touch, and motion during manipulation — high-quality demos for generalization models
Across scenes
One demo, a thousand scene variants
Object-centric representation expands a single demonstration across new poses, layouts, and lighting — no re-collection
Across embodiments
Tabletop skills on a bimanual humanoid
Pick-and-place on a Paxini humanoid shows the same skill transferring zero-shot across robot bodies
Motion retargeting
Map human motion onto a humanoid
SparkUMR retargets everyday actions like drinking onto Fourier GR3 while preserving contact and posture
Fine manipulation
Desktop organization, retargeted
Human motions for arranging a phone and laptop transfer to GR3 — fine manipulation that travels with the skill
Multi-platform
One kitchen skill, many robots
The same mixing and scooping routine drives different humanoid platforms — a unified skill model across embodiments
