Workflow
Promptable Concept Segmentation
icon
Search documents
分割一切并不够,还要3D重建一切,SAM 3D来了
机器之心· 2025-11-20 02:07
Core Insights - Meta has launched significant updates with the introduction of SAM 3D and SAM 3, enhancing the understanding of images in 3D [1][2] Group 1: SAM 3D Overview - SAM 3D is the latest addition to the SAM series, featuring two models that convert static 2D images into detailed 3D reconstructions [2][5] - SAM 3D Objects focuses on object and scene reconstruction, while SAM 3D Body specializes in human shape and pose estimation [5][28] - Meta has made the model weights and inference code for SAM 3D and SAM 3 publicly available [7] Group 2: SAM 3D Objects - SAM 3D Objects introduces a novel technical approach for robust and realistic 3D reconstruction and object pose estimation from a single natural image [11] - The model can generate detailed 3D shapes, textures, and scene layouts from everyday photos, overcoming challenges like small objects and occlusions [12][13] - Meta has annotated nearly 1 million images, generating approximately 3.14 million 3D meshes, leveraging a scalable data engine for efficient data collection [17][22] Group 3: SAM 3D Body - SAM 3D Body addresses the challenge of accurate human 3D pose and shape reconstruction from a single image, even in complex scenarios [28] - The model supports interactive input, allowing users to guide and control predictions for improved accuracy [29] - A high-quality training dataset of around 8 million images was created to enhance the model's performance across various 3D benchmarks [31] Group 4: SAM 3 Capabilities - SAM 3 introduces promptable concept segmentation, enabling the model to identify and segment instances of specific concepts based on text or example images [35] - The architecture of SAM 3 builds on previous AI advancements, utilizing Meta Perception Encoder for enhanced image recognition and object detection [37] - SAM 3 has achieved a twofold improvement in concept segmentation performance compared to existing models, with rapid inference times even for images with numerous detection targets [39]
Meta「分割一切」3.0曝光,技能语义分割加入概念提示,好好玩,要爆了
3 6 Ke· 2025-10-13 03:52
Core Insights - The article discusses the introduction of SAM 3, a third-generation segmentation model that can understand natural language prompts for image and video segmentation tasks [1][3][5]. Group 1: Model Capabilities - SAM 3 can segment images and videos based on user-defined phrases, allowing for more interactive and intuitive segmentation tasks [3][6]. - The model processes images containing over 100 objects in just 30 milliseconds, demonstrating near real-time capabilities for video processing [5][21]. - SAM 3 introduces a new task paradigm called Promptable Concept Segmentation (PCS), which allows for multi-instance segmentation based on various input prompts [6][7]. Group 2: Technical Innovations - The architecture of SAM 3 includes a new detection module based on the Deformable Transformer (DETR), which separates object recognition and localization tasks to enhance detection accuracy [11]. - A scalable data engine was developed to create a training dataset with 4 million unique concept labels and 52 million validated masks, improving the model's performance [12]. - The SA-Co benchmark was introduced to evaluate the model's performance in open vocabulary segmentation tasks, significantly expanding the concept coverage compared to existing benchmarks [13]. Group 3: Performance Metrics - SAM 3 achieved a 47.0% accuracy in zero-shot segmentation tasks on the LVIS dataset, surpassing the previous state-of-the-art (SOTA) of 38.5% [16]. - In the new SA-Co benchmark, SAM 3's performance is at least twice as strong as baseline methods [16]. - The model also outperformed SAM 2 in video segmentation tasks, indicating significant improvements in performance [18]. Group 4: Future Directions - Researchers are exploring the combination of SAM 3 with multimodal large models (MLLM) to tackle more complex segmentation tasks, such as identifying specific scenarios in images [19]. - Despite its advancements, SAM 3 still faces challenges in generalizing to specialized fields like medical imaging and thermal imaging through zero-shot learning [21].