Roboflow
Products
Solutions
Resources
Pricing
Docs
Blog
Toggle theme
Sign In
Playground
Arena
Rankings
Arena Rankings
Vision Evals
Models
Explore
Compare
© 2026 Roboflow
•
Terms
Filter
Playground Demo
By Task
All Models
3D Reconstruction
Captioning
Chart Question Answering
Classification
Depth Estimation
Document Question Answering
Image Embedding
Image Similarity
Image Tagging
Instance Segmentation
Keypoint Detection
Multi-Label Classification
Object Detection
OCR
Open Prompt
Open Vocabulary Object Detection
Phrase Grounding
Pose Estimation
Promptable Concept Segmentation
Promptable Concept Segmentation
Region Proposal
Semantic Segmentation
Video Classification
Video Object Tracking
Vision Language
Visual Question Answering
Zero Shot Segmentation
By Feature
Foundation Vision
LLMs with Vision Capabilities
Multimodal Vision
Real-Time Vision
Zero-shot Detection
By Modality
All
Vision
Multimodal
By Organization
By Access Type
All
Open
Closed
Top Pose Estimation Models
Models that estimate the pose of people or objects in images.
Filters
Sort by:
Newest
YOLOv12
YOLOv12 is an attention-centric real-time object detection model developed by researchers at Tsinghua University, with the arXiv paper published in February 2025 under the AGPL-3.0 license. It introduces an Area Attention module that partitions feature maps into regions and applies self-attention within each region, reducing the quadratic complexity of full self-attention while capturing long-range dependencies. It also incorporates R-ELAN for improved feature aggregation and scaled residual connections for training stability.YOLOv12-L achieves 54.0% AP on COCO, while the YOLOv12-N variant achieves 40.5% mAP at 1.62ms latency on an NVIDIA T4 GPU. The model is built on the Ultralytics codebase, supporting detection, segmentation, and other standard YOLO tasks at competitive real-time speeds.
YOLOv8 Pose Estimation
YOLOv8 Pose Estimation is the keypoint detection variant of the YOLOv8 model developed by Ultralytics, released in April 2023 under the AGPL-3.0 license. It extends the YOLOv8 detection head to predict keypoint locations and visibility scores alongside bounding boxes, using a decoupled head for joint localization and keypoint regression. By default it targets the 17-keypoint COCO human pose skeleton, but can be configured for custom keypoint sets.YOLOv8 Pose shares the same architecture and size variants as the base detection model and achieves competitive performance on the COCO keypoints benchmark at real-time inference speeds. The model is deployable through Roboflow Inference and is suited for applications including sports analytics, ergonomics monitoring, gesture recognition, and human activity detection.
Google
MediaPipe
MediaPipe is an open-source framework developed by Google for building real-time machine learning pipelines across mobile, web, desktop, and edge platforms. First released in 2019, the framework uses a graph-based architecture where pre-built components called Calculators process streaming data such as images, video, and audio through configurable computation graphs. This design allows developers to compose perception pipelines from reusable building blocks without writing custom glue code between models. The current MediaPipe Tasks API replaces the earlier Solutions API and provides a unified cross-platform interface for vision, text, and audio.Rather than providing a single model, MediaPipe ships a suite of ready-to-use Tasks that wrap trained models for specific problems. These include MediaPipe Pose Landmarker for 33-point body landmark detection, Hand Landmarker for 21-point hand tracking, Face Landmarker which extends the earlier 468-point Face Mesh with blendshape outputs for facial expression, Selfie Segmentation for person-background separation, and Holistic Landmarker for combined body, hand, and face tracking. The Tasks prioritize on-device inference with low latency and support GPU acceleration where available, making the framework a common choice for mobile augmented reality, fitness and wellness applications, gesture-based interfaces, and accessibility features such as sign language recognition.