Codec Avatars render a scanned face, not a live scene
Two Meta research papers describe how a photoreal avatar is built and driven, a different claim from a live volumetric feed.
- First source published
- August 14, 2018
- Site publication
- September 18, 2026

What the published papers actually show
Deep Appearance Models for Face Rendering, published at SIGGRAPH in 2018, describes a neural method for rendering a face from a multiview capture setup: a deep model learns geometry and view-dependent texture together from many camera angles of one person's face, then renders new expressions and viewpoints in real time using a standard graphics pipeline. Three years later, Pixel Codec Avatars, an oral presentation at CVPR 2021, describes a lighter version of the same idea built to run several avatars at once, reporting that it can render five avatars in real time on an Oculus Quest 2 for what the paper calls authentic face-to-face communication in three dimensions over remote physical distances. Together the two papers document a lineage, not a single product: a heavy, high-fidelity face model in 2018 and a compressed, multi-person-capable descendant by 2021.
What creates the image, and what it costs to make one
Both systems depend on a per-person training capture: a multiview rig photographs one individual making a wide range of expressions so the model can learn that specific face. That is the mechanism point that matters most. A Codec Avatar is not a camera generating a live three-dimensional scan of whoever sits down, in the way a volumetric capture rig or a depth-camera telepresence system is; it is a personalised, pre-trained model that is then driven, in later headset-based work, by the small inward-facing cameras built into the headset itself. Whoever wears the headset never sees their own avatar's face; what a remote viewer sees is an inference from that pre-trained model, not a live optical image of the wearer at that instant.
Why that is a different claim from a volumetric feed
The distinction is worth holding onto because photoreal avatar and live volumetric capture are often described in the same breath as if they were the same achievement. A volumetric system, such as the multi-camera rigs used in mixed-reality capture studios, reconstructs whoever is standing in the volume, unmodified, from live footage. A Codec Avatar is faster to transmit and can run on consumer hardware precisely because it is not doing that: it is animating a model built in advance from one person's training capture. Neither paper claims otherwise, but the distinction is easy to lose in a demonstration reel.
Questions to bring to a demo
- Is the avatar generated from a live camera feed of the wearer, or driven from a pre-trained personal model.
- How long, and under what capture conditions, did that person's original training session take.
- What happens to fidelity for an expression or angle the training capture never included.
Read this way, the 2018 and 2021 papers are honest about their own scope: they document a face model, trained per person, rendered efficiently. Treating that as equivalent to seeing a whole remote body live is the reader's error, not a claim either paper makes.
Sources & reading trail
Describes the multiview capture and deep variational model behind photoreal facial rendering, presented at SIGGRAPH 2018.
Source published: 14 August 2018 · Retrieved: 16 September 2026
Describes a lighter, multi-person Codec Avatar model rendering five avatars in real time on Oculus Quest 2, presented at CVPR 2021.
Source published: 9 April 2021 · Retrieved: 16 September 2026
Primary documents establish the record; the mechanism reading and the demo questions are Presence Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.
Continue reading
- MetaHuman Creator ships digital humans as reusable rigs
- NeRF learned a scene's geometry from ordinary photographs
- A scanned Persona stands in for a face a headset hides
- Browse the complete the archive
Sources & reading trail
- Deep Appearance Models for Face Rendering
Source published: August 14, 2018 · Retrieved: September 16, 2026 - Pixel Codec Avatars
Source published: April 9, 2021 · Retrieved: September 16, 2026
The documents above establish the record. The reading and the questions are this publication’s editorial analysis, written after the fact.
Published September 18, 2026, not on the date of the event described.