Abstract
Augmented Reality (AR) telepresence has emerged with the advance of immersive technologies such as Head-Mounted Displays (HMDs). It enables remotely separated users to meet with each other through virtual agents merged in their own spaces as if they are face-to-face. Compared to conventional video conferencing, it could dramatically increase the sense of being spatially co-present with remote peers. However, such increased spatiality introduces more factors to concern regarding each agent’s relation with the telepresence space, the local user, and other telepresence agents. Moreover, the intrinsic delay sensitivity and the high mobility and accessibility requirements of future lightweight always-on AR HMDs add further constraints to the application design. First, users’ telepresence spaces are often different in size and layout, making it necessary to adapt agents’ motions according to the physical scene. How to achieve semantically right spatial-temporal adaptations with low delays in a room-scale experience is under-explored. Second, users’ perceived social proxemics in AR experiences tend to be different from those in the physical world due to the limitations of AR HMDs, such as their small Field of View (FoV). Studies on the quantitative spatial adaptation of user-centered telepresence agent placement regarding interpersonal distance and the sense of co-presence are still rare. Third, spatial-faithful attention cues indicating “who is looking at whom” are often difficult to present in multiparty AR telepresence, especially in solutions not based on 3D avatars. There exist few feasible solutions to adapting agents’ attention cues in HMD-based AR telepresence with an accessible setup. Targeting the three problems introduced above, this thesis explores spatial, temporal, and attentive adaptation for AR telepresence agents and presents novel techniques to adapt the agent’s motion, placement, and attention cues.We first study spatial-temporal adaptation for agent motion regarding the telepresence space. Different telepresence spaces have heterogeneous structures and features, which bring difficulties in synchronizing telepresence avatar motions with real user motions and adapting avatar motions to local scenes. Existing methods generate mutual movable spaces or retarget the placement of avatars. However, they limit the telepresence experience in a small sub-area or present delayed and discontinuous avatar transitions. Focusing on a single-avatar scenario, we first examine the impact of the transition delay and explore the transition style with such delay through user studies. With the obtained design, we propose a Predict-and-Drive controller based on object-to-object mapping and an artificial potential field to diminish the delay and present smooth avatar transitions. We further conduct ablation studies to evaluate the effectiveness of our proposed components.
After addressing the motion adaptation regarding a single agent’s interaction with the space, we study the spatial placement adaptation regarding the social distance between remote agents and the local user in video-avatar-based multiparty AR telepresence. Real-time volumetric AR telepresence methods are often too complex for everyday usage. Other solutions target mobile and effortless-to-setup telepresence on AR HMDs, but can only support either high fidelity or a high level of co-presence. To achieve a balance between fidelity and co-presence, we explore using life-size 2D video-based avatars. We first conduct a pilot study to explore the optimal video avatar placement in AR conversations. With the placement results, we then implement a proof-of-concept prototype and conduct user evaluations to verify its effectiveness in balancing fidelity and co-presence. We further quantitatively model the effect of FoV size on video avatar layout to adapt telepresence avatar placement for various HMDs.
Besides the spatial-temporal adaptation with respect to the local user and space, we further explore agents’ attention adaptation in multiparty AR telepresence. To make users aware of non-verbal cues indicating “who is looking at whom”, existing solutions explore screen-based visualizations, incorporate additional hardware, or alter to use a virtual avatar representation. However, they lack immersion, are too complex for everyday usage, or lose the fidelity of remote users’ appearances. In our approach, we conduct two user studies to propose an “Attention Circle” layout and a “rotatable 2.5D video avatar with attention thumbnail” visualization to achieve full attention awareness in multiparty AR telepresence. With the obtained design, we implement A3RT, a proof-of-concept prototype system that empowers attention-aware 2.5D-video-avatar-based multiparty AR telepresence in an everyday setup. Ablation and usability studies on the prototype verify the effectiveness of the proposed components and the full system.
| Date of Award | 18 Jun 2024 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Hongbo FU (Supervisor), Weizhan ZHANG (External Supervisor) & Christian SANDOR (External Co-Supervisor) |
Cite this
- Standard