Abstract:3D animatable human avatars are widely used in film making, game characters, and telepresence. To create these avatars efficiently, researchers propose to reconstruct the avatar from several input images with a feedforward network. These networks are mostly transformers trained with large proprietary data and thousands of H100 GPU hours. However, do we really need such a scale of data and compute for this task?
Our answer is no. In this talk, we present ARG-Avatar, a lightweight network with only 68M trainable parameters, yet achieves SOTA performance on OOD testing data with 13x less training compute to the best baseline. We will go through the two core components of ARG-Avatar. The first is FACRoPE, which inject geometric guidance into the attention with RoPE mechanism. We propose a novel coordinate formulation named Foreground Avatar Coordinates (FAC), to associate let the network tokens focus on its corresponding image regions. The second is Intermediate Token Rendering (ITR), which decodes a coarse avatar from intermediate network tokens during forward pass. We show that these components are all beneficial in the ablation study. In conclusion, we show that by properly injecting geometric into attention, we make an architecture that learns more efficiently and effectively for feedforward avatar reconstruction.
—
