Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>In recent years, human pose estimation has made significant progress in computer vision. However, existing methods still face two major challenges. First, heatmap-based deterministic representations are insufficient for modeling the uncertainty of keypoint predictions and cannot effectively handle out-of-image keypoints. Second, single-scale feature modeling fails to capture both local details and global structural information, resulting in limited performance under occlusion and truncation scenarios.To address these issues, we propose a probabilistic modeling-based framework for 2D human pose estimation. Specifically, we build upon the ProbPose framework and introduce a Multi-Scale Context Fusion Vision Transformer (MSCF-ViT) as the backbone to enable effective fusion of multi-scale features through cross-scale interactions. Furthermore, an Adaptive Probability Map (APM) module is designed to dynamically adjust spatial probability distributions based on keypoint types and occlusion conditions, enhancing uncertainty modeling capability. To further improve performance, we propose an Adaptive OKSLoss (A-OKSLoss) that aligns the training objective with the evaluation metric while strengthening the learning of hard samples.Extensive experiments on COCO2017, CropCOCO, and OCHuman datasets demonstrate that the proposed method outperforms existing approaches, especially in challenging scenarios involving occlusion and out-of-bound keypoints.</p>

Show More

Keywords

modeling occlusion human pose estimation

Related Articles

PORE

About

Connect