SenseTime has announced the open-source lightweight unified multimodal model SenseNova U1.5-Lite-Preview for communities. This model is not a simple scale upgrade, but a systematic iteration based on SenseNova U1, integrating four key capabilities—visual understanding, reasoning, generation, and editing—within the same architecture.
The most notable feature is its lightweight size—achieving 4K resolution, more refined realistic textures, and more complex visual control with only an 8B-MoE parameter scale. SenseTime stated that if U1 validated the feasibility of the unified architecture, then U1.5 aims to further prove whether the capabilities can continue to expand.

Significant improvements in four key areas, full-chain optimization of generation and editing
Compared to U1, U1.5-Lite-Preview brings significant improvements in four core areas: native support for 4K image generation, more refined local textures and realistic world textures, more accurate Chinese and English text generation and complex layout organization, and more stable image editing and visual instruction following.
In technical details, SenseTime redesigned the generation head to address grid artifacts and high-resolution detail modeling issues, reducing the impact of visual tokens on the final image and expanding training to 4K resolution. In image editing, the model reorganized editing data and training tasks, enhancing the ability to maintain original content, subject identity, spatial structure, and non-editing areas, ensuring that while achieving the desired modification, as much as possible of the parts not needing changes in the original image are preserved.
Significant improvements in multiple benchmark tests, especially in editing capabilities
Evaluation results confirm the actual effects of this optimization. Qwen-Image-Bench scores increased from 47.14 to 55.20 (with prompt enhancement), ImgEdit-Bench from 3.90 to 4.37, GEdit-Bench English version from 7.47 to 8.17, and Chinese version from 7.42 to 8.05. Overall, both generation quality and editing stability have seen significant improvements, particularly in instruction following and maintaining the original image during image editing.
SenseTime emphasized that the core idea of U1.5-Lite-Preview is not to blindly expand the model scale, but to systematically optimize the entire generation and editing process. This approach of "gaining capability growth through architectural innovation" provides a noteworthy model for the continuous iteration of lightweight multimodal models—when 8B parameters can run 4K generation and fine editing, the potential of the unified multimodal architecture may just be beginning to unfold.
