Objects, Scenes, and People: from Observation to Reconstruction
Abstract
Recovering the three-dimensional structure of objects, scenes, and people from images requires reasoning beyond what is directly visible---behind occluding surfaces, outside the field of view, and across time. This thesis develops methods for reconstructing shapes and motion across these domains, from static objects and scenes to human pose in dynamic settings. A recurring theme is fusing complementary representations: coordinate frame choices, perceptual and generative approaches, egocentric and exocentric perspectives. We begin by studying how coordinate systems and viewpoint choices shape learned representations for object shape prediction and pose estimation. We then develop multi-layer depth representations as a viewer-centered alternative to detection-based scene reconstruction. Combining viewpoint prediction with pose estimation produces representations that transfer across datasets. Building on these foundations, we integrate visual tracking with generative motion modeling to recover global multi-person motion by fusing egocentric and exocentric views.
Cite
@phdthesis{objects-scenes-and-people-from-observation-to-reconstruction-2026,
author = {Daeyun Shin},
title = {Objects, Scenes, and People: from Observation to Reconstruction},
school = {University of California, Irvine},
year = {2026},
url = {https://escholarship.org/uc/item/9xg0v169},
}