CivArchive
    SenseNova-U1.5-8B-MoT - SFT
    NSFW
    Preview 141504414
    Preview 141504406
    Preview 141504390
    Preview 141504394
    Preview 141504404
    Preview 141504405
    Preview 141504410
    Preview 141504413

    This model is not mine. See official repo at https://huggingface.co/sensenova/SenseNova-U1.5-8B-MoT

    The following is the original README (minus diffusers installation stuff)


    SenseNova-U1.5 native unified multimodal architecture

    Overview

    SenseNova-U1.5-8B-MoT is our latest native unified multimodal checkpoint for more accurate, consistent, reliable, and aesthetically compelling visual creation. Built on NEO-unify, it strengthens the patchify layers, data quality and distribution, task formulation, prompt enhancement, and post-training pipeline.

    The official release focuses on six user-visible improvements:

    • Higher-quality image generation: improved composition and color harmony, with more realistic material rendering, natural lighting, stronger visual fidelity, and finer local details.

    • Better text rendering and infographic generation: more legible Chinese and English text, with clearer information hierarchy in posters, infographics, brand assets, and other text-dense designs.

    • More efficient native 4K generation: more coherent global structure, color harmony, and stable high-resolution output with improved generation efficiency.

    • More reliable native image editing: stronger preservation of subject identity and unedited content across local, text, multi-reference, insertion, and replacement edits.

    • Stronger complex-instruction following: more consistent execution of object counts, spatial relationships, layouts, styles, and multiple constraints within a single request.

    • More precise visual control: more accurate region- and object-level control through bounding boxes, visual markers, and single- or multi-image references.

    Showcases

    SenseNova-U1.5 generation and editing showcases

    Key Benchmarks

    SenseNova-U1.5 benchmark overview

    View detailed benchmark results

    SenseNova-U1.5 detailed benchmark results

    Best Practices

    Direct natural-language prompts work well for clear tasks with few constraints. For complex generation or editing, use prompt enhancement when additional planning is needed and explicitly specify what should remain unchanged.

    See the SenseNova-U1.5 Cookbook for setup instructions and optional Image PE, Caption-to-Prompt, and Editing PE recipes.

    ๐ŸŒ Use with SenseNova-Studio

    The fastest way to experience SenseNova-U1.5 is through SenseNova-Studio โ€” a ๐Ÿ†“ free online playground where you can try the model directly in your browser, no installation or GPU required.

    Ongoing Improvements

    The official release improves upon the Preview, though challenges remain in:

    • Over-emphasized details or colors: some prompts may produce excessive high-frequency detail or oversaturated colors, which can often be mitigated by lowering cfg_scale.

    • Dense text errors: dense, lengthy, small, or mixed Chinese-English text may contain errors.

    • Constrained layouts: exact counts, alignment, or hierarchy may be imperfect in highly constrained layouts.

    • Unstable human details: small faces, hands, limbs, and fine-grained object structures may remain unstable.

    • Complex editing drift: broad, multi-turn, or multi-reference edits may drift, especially when many regions must be preserved simultaneously.

    Models

    ModelStageHF WeightsSenseNova-U1.5-8B-MoTRL๐Ÿค— ModelSenseNova-U1.5-8B-MoT-SFTSupervised fine-tuning๐Ÿค— Model

    ๐ŸŒ Join the Community!

    Join our growing community to share feedback, get support, and stay updated on the latest SenseNova-U1 developments โ€” we'd love to hear from you!

    DiscordFeishu Group

    Citation

    If this project is helpful for your research, please consider starring the repository and citing:

    @misc{sensenova2026neounify,
      title        = {NEO-unify: Building Native Multimodal Unified Models End to End},
      author       = {SenseNova},
      journal      = {Hugging Face blog},
      url          = {https://huggingface.co/blog/sensenova/neo-unify},
      year         = {2026}
    }
    
    @article{sensenova2026sensenovau1,
      title        = {SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture},
      author       = {Diao, Haiwen and Wu, Penghao and Deng, Hanming and Wang, Jiahao and Bai, Shihao and Wu, Silei and Fan, Weichen and Ye, Wenjie and Tong, Wenwen and Fan, Xiangyu and others},
      journal      = {arXiv preprint arXiv:2605.12500},
      year         = {2026}
    }
    

    License

    This model is released under the Apache 2.0 License.

    Description

    SFT Version

    Comments (4)

    sriramakumar7799888Sep 1, 2026ยท 1 reaction
    CivitAI

    Looks interesting. Could you please share the text encoder and VAE details, along with a workflow.

    clueless_engineer
    Author
    Sep 1, 2026ยท 2 reactions
    sriramakumar7799888Sep 1, 2026ยท 2 reactions

    @clueless_engineerย Thank you.

    aj_prime8517Sep 2, 2026

    I have been using it for last 10 days, it is nice for making comic pages
    https://civitai.red/images/141017237
    https://civitai.red/images/140912728
    but still not perfect when a lot of text is used
    https://civitai.red/images/140950237
    here some edit
    https://civitai.red/posts/30618356
    and here I used with prompts that I already used with other models
    https://civitai.red/posts/30565219

    Checkpoint
    Other

    Details

    Downloads
    37
    Platform
    CivitAI
    Platform Status
    Available
    Created
    9/1/2026
    Updated
    9/7/2026
    Deleted
    -

    Files

    sensenovaU158BMot_sft.safetensors

    sensenovaU158BMot_sft.safetensors

    Mirrors