SenseTime has announced the open-source preview of its lightweight, unified multimodal model SenseNova U1.5-Lite-Preview. As a systematic iteration of SenseNova U1, this model is based on the NEO-unify architecture, achieving native 4K image generation, more detailed texture expression, more accurate Chinese and English text generation, and more stable image editing capabilities with only an 8B-MoT parameter scale.

QQ20260803-173044.jpg

According to the introduction, the SenseNova U1.5-Lite-Preview significantly improves the model's ability to understand long natural language descriptions and structured visual instructions, enabling more precise visual control according to complex requirements. In terms of image editing, the new model supports style transfer from reference images, combined editing with multiple reference images, and interactive precise adjustment, allowing AI image creation to shift from one-time generation to continuous iteration and optimization.

Compared to the previous generation model, U1.5-Lite-Preview has achieved improvements in multiple evaluation benchmarks. SenseTime stated that this model can balance generation quality and interaction capabilities with a smaller model size, further lowering the entry barrier for the application of multimodal AI technology.

Currently, SenseNova U1.5-Lite-Preview is available for download on GitHub, Hugging Face, and ModelScope community, allowing developers and researchers to perform secondary development based on the model. At the same time, the delivery-level creative model for professional creators and enterprise users SenseNova U1Pro is in a tight invitation testing phase.

As the competition in multimodal models gradually shifts from parameter scale to efficiency, control capability, and practical application experience, SenseTime's open-source of this lightweight model also reflects the industry's trend of promoting high-performance AI capabilities toward a broader developer ecosystem.