CivArchive
    ← All articles
    Published August 8, 2026by liutyi

    SeFi Image

    110 views2 reactions1 comments on CivitAI0 collected
    comparative study

    Intro

    The SeFi is family of 7 small models. Published ~ 17.06.2026.

    Se and Fi stands for Semantic First.

    1. SeFi-Image-1B-turbo (4 STEPS, CFG 1)

    2. SeFi-Image-2B-turbo (4 STEPS, CFG 1)

    3. SeFi-Image-5B-turbo (4 STEPS, CFG 1)

    4. SeFi-Image-1B-Base (50 STEPS, CFG 4)

    5. SeFi-Image-2B-Base (50 STEPS, CFG 4)

    6. SeFi-Image-5B-Base (50 STEPS, CFG 4)

    7. SeFi-Image-5B-RL (50 STEPS, CFG 4)

    HuggingFace repos are gated. Your must be logged in and accept that the models are for non-commercial use only. License is CC BY-NC 4.0.

    Base vs Turbo

    Base and RL are much better than Turbo

    Latest models got random approach to Base vs Turbo image quality. Some model families got both base/turbo usable (ERNIE, Z Image, Boogu) . Some leave base almost unusable and Turbo is nice and shiny (like KREA2). And some do the opposite (Mage Flow) - base models usable and turbo is not. SeFi is in the last group. Where Turbo is done wrong. And better to use Base (slow) models.

    Compare model size

    To get some understanding of how small this model we will compare it to other models. Since usual "Billions of parameters" is somehow misleading (like, for example, Microsoft Lens 4B is actually 24B in total parameters) we will compare Total parameters:

    • 3.3B - SeFi Image 1B

    • 3.4B - Stable Diffusion XL (Pony, Illustrious)

    • 4.3B - SeFi Image 2B

    • 6.9B - Cosmos Predict 2B (Anima)

    • 8.6B - Microsoft Mage Flow 4B

    • 9.4B - SeFi Image 5B

    • 10.2B - Z Image 9B

    • 14.6B - Long Cat Image 6B

    • 15.3B - ERNIE Image 8B

    • 17.3B - FLUX.2 Klein 9B

    • 17.4B - FLUX.1 Dev 12B

    • 16.8B - KREA 2 Turbo 12B

    • 18.5B - Boogu Image 10B

    • 25.1B - Microsoft Lens 4B

    • 26.3B - Hunyuan Image 2.1 17B

    • 26.8B - Ideogram 4 9B

    • 28.8B - Qwen Image 20B (Qwen Image 2512)

    • 30.8B - HiDream I1 Full 17B

    • 56.3B - FLUX.2 Dev 32B

    Models on CivitAi

    Officially not present. The unofficial are:

    Image Quality

    We are not expecting high image quality from 1B/2B models, are we? But for 5B it is some hope. I cannot remember any <8B models that are generally good, but maybe this time? Since it very close to Z Image by total parameters count, why not? So let us check..

    SeFi Image 5B RL

    (from test v1). CFG 4, STEPS 50, first attempt image on seed 20260804

    RL produce results close to Base

    SeFi Image 5B Base

    (from test v1). CFG 4, STEPS 50, first attempt image on seed 20260804

    so it can draw text, it can do some photorealism, ok, how about some memes?

    it is definitely not perfect, but might be worse, and it actually is on lower sizes

    SeFi Image 2B Base

    it still work with text (being smaller than Cosmos Predict 2B/Anima). So what about 1B?

    SeFi Image 1B Base

    it is bad. but wait a second.. it still able to position and render some text. being less than SDXL size in total parameters.. Impressive. Now Let's switch to Turbo models and that is when all starts to be ugly

    SeFi Image 5B Turbo

    SeFi Image 2B Turbo

    SeFi Image 1B Turbo

    4/8/10 Steps

    increasing steps for Turbo is not help much

    (from test v3) "Woman laying on a grass" classic

    1B Base 4 STEPs

    1B Base 8 STEPs

    1B Base 10 STEPs

    5B RL, 50 steps to compare

    Text in 1B

    What is most impressive? I guess 4 steps 1B model that able to create image with correct text like

    and that is fast. that is like 6 seconds on Intel iGPU (285H). not all generations got that image correct, but from 2-5 attempt, something like this might be result of prompt:

    Graphic design poster with a tactile mixed-media aesthetic. The exact main word “SeFi” is centered and dominant, spelled correctly with four clearly readable characters: "S", "e", "F", "i". The letters are thick, rounded, three-dimensional forms made from realistic blue modeling clay, with visible fingerprints, subtle cracks, soft irregular edges, and natural clay texture. The word “SeFi” must remain perfectly legible and must not be altered, misspelled, merged, duplicated, or replaced by pseudo-text. Directly underneath, centered below “SeFi,” is the exact smaller word “Image”, clearly spelled "I", "m", "a", "g", "e", printed in dark black typography on a narrow torn rectangular fragment of an old newspaper page. The newspaper fragment has visible paper fibers and tiny authentic newspaper print surrounding the large readable word “Image”. In the upper-left corner, place a small bronze hexagonal badge containing the exact characters “1B”, cleanly rendered and highly legible, with embossed metallic lettering and subtle aged bronze patina.
    The entire composition sits on textured paper with a restrained gradient transitioning from dark charcoal grey on the left and center into deep muted dark red toward the right and lower areas, with grey strongly dominating the overall palette. The paper surface contains sparse abstract geometric construction marks: thin circles, hexagons, angular lines, measurement-like marks, faded diagrams, scratches, pencil strokes, and distressed print imperfections. Keep these background elements subordinate to the typography. Strong studio lighting creates realistic shadows beneath the clay letters, newspaper fragment, and bronze badge, giving the poster physical depth while maintaining a clean graphic composition. High-quality editorial graphic design, tactile collage, realistic materials, crisp edges, controlled contrast, centered typography, generous negative space, extremely accurate text rendering. The three required text elements must be exactly: “SeFi”, “Image”, and “1B”.

    Links for Test results

    Test v1

    Test v2

    Test v3

    Ai Art

    RL

    Turbo

    Archived from CivitAI · Updated August 14, 2026View source