A humanoid android modeled on actress Cai Ming stunned audiences during the 2026 Spring Festival Gala, appearing almost indistinguishable from a real person. Behind such increasingly lifelike digital humans lies a technological breakthrough now emerging from China: the construction of the country’s largest high-precision three-dimensional facial database, designed to dramatically improve how machines perceive and reproduce human faces.
The new database, developed by a research team led by Professor Song Zhan at the Shenzhen Institute of Advanced Technology under the Chinese Academy of Sciences, in collaboration with Dr. Ye Yuping from Fujian University of Technology, contains approximately 200,000 high-fidelity 3D facial scans. The work, published in IEEE Transactions on Circuits and Systems for Video Technology, represents one of the most comprehensive efforts to address a longstanding bottleneck in 3D facial keypoint detection.
Three-dimensional facial keypoint detection is a core technology enabling digital humans and humanoid robots to display vivid emotions, recognize identities and perform embodied interactions. It allows machines to locate precise anatomical landmarks—such as the corners of the eyes, the tip of the nose and the contours of the mouth—in three-dimensional space. These landmarks are essential for realistic facial animation, biometric authentication and human–machine interaction.
Until now, however, progress in the field has been constrained by the lack of large-scale, accurately annotated 3D facial datasets. Most existing algorithms rely heavily on two-dimensional texture mapping or artificially generated digital 3D faces. Such approaches are limited by texture inaccuracies and by the gap between synthetic models and real human facial geometry, leading to reduced performance in real-world scenarios.
To overcome these limitations, the research team built a custom 3D and 4D facial acquisition system and carried out standardized data collection procedures. The resulting multimodal biometric database goes beyond static scans. In addition to the 200,000 high-fidelity 3D facial models, it includes a multi-expression 3D face dataset, a standardized 3D facial landmark dataset, a high-precision 3D human body dataset and a dynamic 4D facial expression dataset that captures changes over time.
The scale and quality of the dataset have already drawn official recognition. It was selected for Fujian Province’s 2025 High-Quality AI Dataset Program, highlighting its strategic importance for artificial intelligence development.
Alongside the dataset, the team introduced a new algorithm known as a curvature-fused graph attention network, or CF-GAT. Unlike traditional methods that depend on ordered image grids or pre-defined templates, CF-GAT operates directly on raw point clouds—unordered collections of 3D data points representing the geometry of a face.
The researchers adopted a geometry-driven sampling strategy to simplify point clouds while preserving essential curvature information. This curvature data was encoded as an explicit geometric prior and integrated into the model’s attention mechanism. By doing so, the network can focus on subtle local shape variations that define facial features, while also modeling broader global relationships among points.
Through its graph attention structure, CF-GAT predicts 3D landmark coordinates directly from geometric data, eliminating the need for 2D texture assistance or template-based alignment. This shift allows the system to learn richer geometric patterns and adapt more effectively to real-world facial diversity.
Experimental results indicate that the network demonstrates stronger robustness to noise, improved generalization across varied facial shapes and more precise localization of fine-grained landmarks. In practical terms, this means better performance in complex environments where lighting conditions, occlusions or facial variations might otherwise degrade accuracy.
The implications extend far beyond entertainment. More accurate 3D facial landmark detection could enhance humanoid robotics, allowing androids to replicate nuanced human expressions with greater realism. It could also improve virtual avatars in gaming and online communication, strengthen biometric security systems and support medical or psychological research that relies on detailed facial analysis.
The dynamic 4D component of the dataset, capturing temporal changes in expression, may be particularly valuable for developing embodied intelligence—systems capable not only of recognizing faces but also of interpreting and responding to emotional cues in real time.

