Workflow
多模态交互技术
icon
Search documents
商汤科技李星冶:多模态大模型“所见即所得”让人机交互更顺畅
Bei Ke Cai Jing· 2025-07-10 11:49
Core Insights - The article discusses the evolution of artificial intelligence from 1.0 to 2.0, highlighting SenseTime's breakthroughs in multimodal interaction technology and its applications across various sectors [1][2]. Group 1: AI Evolution - SenseTime has transitioned from focusing on computer vision in the AI 1.0 era to promoting multimodal interaction innovations in the AI 2.0 era, driven by the rise of large model technologies in 2023 [1]. - The concept of "seeing is believing" is emphasized, integrating video, images, and voice to enable real-time interaction with humans [1]. Group 2: Applications in Education - In the education sector, SenseTime collaborates with learning device manufacturers to develop interactive devices that utilize real-time algorithms to assist children in solving problems and recognizing errors [2]. - The system supports interactive storytelling for young children by converting images into narratives, and SenseTime has partnered with around 10 schools to create smart campus assistants for managing course schedules and grade inquiries [2]. Group 3: Intelligent Applications - SenseTime's intelligent applications include algorithms that analyze industry data to assist in warehouse leasing scenarios and generate lease management solutions [2]. - In customer service, SenseTime collaborates with well-known operators to create efficient intelligent agents, and in smart home applications, it enhances family interaction through AI technology [2]. - The advantage of multimodal large models lies in enabling smoother interactions beyond text command recognition, utilizing visual and multidimensional information [2].