Abstract
3D digital humans play crucial roles in a wide range of applications, including autonomous driving, live streaming, virtual/augmented reality, film production, and human–computer interaction. As these applications continue to evolve, the demand for realistic and scalable digital human creation has become increasingly urgent. Although multi-view capture systems or specialized rigs can produce high-quality 3D humans, they are expensive, labor-intensive, and impractical for large-scale deployment. This limitation has motivated growing interest in reconstructing digital humans directly from monocular RGB images. However, recovering high-fidelity geometry and appearance from only a single view remains extremely challenging due to severe self-occlusions, diverse clothing, complex illumination, and the requirement for consistent modeling of both geometry and appearance.Despite these difficulties, numerous attempts have been made to reconstruct 3D humans from monocular images. Current approaches can be broadly grouped based on their 3D representation into implicit and explicit methods. Implicit-based methods aim to recover continuous geometry and appearance fields, but they often experience feature ambiguity and lack sufficient supervision for producing realistic textures. Many explicit-based methods, which typically represent humans with meshes or other discrete representations, alleviate occlusions by first generating additional viewpoints with large 2D image diffusion models. However, the synthesized views often contain positional misalignments and missing details, which further hinder high-fidelity reconstruction.
In this thesis, a series of explicit diffusion-based 3D human reconstruction methods is introduced to overcome the limitations of the two prevailing paradigms in 3D digital human reconstruction. The first method, HaP, employs point cloud diffusion to recover high-quality human geometry from a single RGB image. Although HaP focuses on geometric reconstruction, its generated point clouds provide accurate and stable 3D anchors that naturally align with the centers used in 3D Gaussian Splatting. This observation motivates HuGDiffusion, which takes HaP’s geometry as initialization and learns only the remaining appearance-related attributes—such as spherical harmonics, scales, opacities, and rotations—to achieve high-quality novel view synthesis. While HaP and HuGDiffusion address geometry recovery and appearance modeling, respectively, they still treat the two components separately. Such decoupled designs are common in existing pipelines but often lead to inefficiency and inconsistencies between geometry and appearance. To resolve this fundamental limitation, this thesis further proposes JGA-LBD, a unified framework that jointly models both modalities. JGA-LBD employs a sparse variational autoencoder to encode geometry and appearance into a shared latent space, and a bridge diffusion model to generate unified latent representations that capture both aspects simultaneously. With a single decoding stage, the framework reconstructs detailed 3D meshes and produces coherent novel views, achieving efficient and consistent digital human reconstruction.
We first introduce HaP, a point cloud diffusion-based reconstruction framework that differs from implicit methods by explicitly generating and manipulating 3D point clouds in 3D space. HaP begins by estimating depth maps and SMPL models from input RGB images. The depth maps are converted into partial point clouds, which are used to rectify the pose and shape of the SMPL estimates. The combined partial point cloud and SMPL geometry serve as conditioning inputs for training a point cloud diffusion model that generates complete human shapes. A subsequent refinement module and a surface reconstruction module further enhance geometric fidelity and robustness.
Building upon HaP, HuGDiffusion extends the reconstruction pipeline from geometry to novel view synthesis. It formulates single-image human novel view rendering as a conditional diffusion process over 3D Gaussian attributes, establishing a generalizable human 3D Gaussian Splatting framework. To enable effective supervision, a two-stage ground truth 3D Gaussian attributes preparation strategy is introduced. In addition, human-centric conditioning signals—such as pixel-aligned features and SMPL semantic features—are incorporated to improve visual realism and generative stability.
Finally, JGA-LBD is proposed to unify geometry reconstruction and appearance modeling within a single framework. By converting heterogeneous inputs—depth maps and SMPL models—into a consistent 3D Gaussian representation and compressing them with a sparse variational autoencoder, JGA-LBD enables joint modeling of geometry and appearance through a bridge diffusion process. This unified design allows the system to reconstruct high-resolution 3D humans and produce photorealistic novel views in a single decoding stage.
Collectively, this thesis establishes a progressive framework for single-image digital human reconstruction, evolving from explicit 3D point–based modeling to diffusion-driven appearance generation and finally to unified latent-space reconstruction. HaP, HuGDiffusion, and JGA-LBD form a consistent trajectory toward integrating geometry and appearance, improving reconstruction fidelity while enhancing generalizability and computational efficiency. Together, these contributions move digital human reconstruction closer to practical, scalable, and unified single-image 3D digitization.
| Date of Award | 2 Mar 2026 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Junhui HOU (Supervisor) |
Cite this
- Standard