This paper studies BMI, weight, and height estimation from a single human image captured in the wild. The authors introduce a dataset of 6,105 images with ground-truth height, weight, and BMI labels, covering varied ethnicities, ages, genders, poses, backgrounds, image scales, and partial occlusions. They compare RGB, depth-map, pose-affinity-map, and edge-map modalities with single-task and multi-task learning, using VGG, DenseNet, and ResNet backbones. According to the supplied abstract, full-body images perform better than half-body and face-only images, although detailed metrics and evaluation protocols are not provided.
No heat snapshots are available in the last 24 hours.