SenseTime Releases Two Multimodal Foundation Models, Driving Visual AI from 'Building Blocks' to Native Unification
Around the 2026 World Artificial Intelligence Conference (WAIC), SenseTime released two multimodal foundation models: the delivery-level agent foundation model SenseNova U1 Pro for long-horizon tasks, and the open-source vision model SenseNova-Vision. The former is based on the NEO-unify native unified architecture, integrating understanding, generation, and action. It supports 8K native ultra-high-definition output, interleaved text-image reasoning, and long-horizon agentic loops, enabling end-to-end tasks from information gathering to visual delivery. The latter unifies classic vision tasks such as object detection, segmentation, and depth estimation as multimodal generation problems, discarding traditional task-specific heads and performing end-to-end modeling in a shared representation space. It achieves state-of-the-art (SOTA) results in structured understanding while approaching the performance of top expert models. These two models address upper-level visual creation and delivery and lower-level physical world perception, respectively, advancing SenseTime's long-term goal of building a unified 'full-perception, full-generation' multimodal foundation model.
Also available in 中文.