中文
← Back to news
ModelsJul 21, 2026

SenseTime Releases Two Multimodal Foundation Models, Driving Visual AI from 'Building Blocks' to Native Unification

Around the 2026 World Artificial Intelligence Conference (WAIC), SenseTime released two multimodal foundation models: the delivery-level agent foundation model SenseNova U1 Pro for long-horizon tasks, and the open-source vision model SenseNova-Vision. The former is based on the NEO-unify native unified architecture, integrating understanding, generation, and action. It supports 8K native ultra-high-definition output, interleaved text-image reasoning, and long-horizon agentic loops, enabling end-to-end tasks from information gathering to visual delivery. The latter unifies classic vision tasks such as object detection, segmentation, and depth estimation as multimodal generation problems, discarding traditional task-specific heads and performing end-to-end modeling in a shared representation space. It achieves state-of-the-art (SOTA) results in structured understanding while approaching the performance of top expert models. These two models address upper-level visual creation and delivery and lower-level physical world perception, respectively, advancing SenseTime's long-term goal of building a unified 'full-perception, full-generation' multimodal foundation model.

Also available in 中文.