🎨 AniSee
A personal anime model built on Anima.
Two generations · Five downloads · Tags + natural language · ComfyUI-ready · LoRA-friendly
AniSee started as a full fine-tune of Anima Preview3 Base and has since moved onto the final Anima v1.0 / v1.1 releases. Both generations stay available, because they are genuinely different things and some people will prefer the older one.
📦 The Five Downloads
⭐ AniSee-V2-Turbo — start here
Built on anima-turbo-v1.1 · LoKr merge · ~3.9 GB
12 steps, CFG 1 — around 6 seconds per image
Diffusion model only — text encoder and VAE needed separately
🎨 AniSee-V2-Aesthetic — the best-looking default
Built on anima-aesthetic-v1.1 · LoKr merge · ~3.9 GB
40 steps, CFG 4–5 — richer colour, softer light
Diffusion model only — text encoder and VAE needed separately
🧪 AniSee-V2-Base — the flexible one
Built on anima-base-v1.0 · LoKr merge · ~3.9 GB
40 steps, CFG 4–5 — most variety, most neutral style
The right base if you want to train your own LoRAs on AniSee
Diffusion model only — text encoder and VAE needed separately
📦 AniSee AIO-V1 — the easy one
Built on Anima Preview3 Base · full fine-tune · ~5.5 GB
One single file. Image model, Qwen text encoder and Qwen-Image VAE all inside
Loads with the standard ComfyUI Checkpoint Loader — nothing else to install
Still the only all-in-one build
🟢 AniSee v1 — the original
Built on Anima Preview3 Base · full fine-tune · ~5.2 GB
Diffusion model variant, drop-in for
anima-preview3-base.safetensors
🤔 Which one should I download?
Just want to generate, fast? → AniSee-V2-Turbo
Want the nicest look out of the box? → AniSee-V2-Aesthetic
Want maximum variety, or want to train LoRAs? → AniSee-V2-Base
Want one file and zero setup? → AniSee AIO-V1
Already running the classic Anima Preview3 setup? → AniSee v1
🆕 What Changed in V2
V2 moves from Anima Preview3 onto the final Anima v1.0 / v1.1 releases, and it changes how the model was made.
V1 was a full fine-tune. V2 is not. V2 is a LoKr merge: a LoKr network — linear 64 / alpha 64, conv 16 / alpha 16, full-rank, factor 4 — was trained for roughly 24,000 steps on Anima-Base-v1.0, then baked into three different Anima variants at strength 1.0. That is a lighter touch than a full fine-tune, and it is the honest description of what these files are.
The upside of working that way: one training run gives three flavours. Same AniSee character, three foundations with genuinely different behaviour.
⚠️ V2 is not an all-in-one checkpoint
This is the one thing worth reading before downloading. Unlike AIO-V1, the V2 files are plain diffusion model files. Text encoder and VAE are not included. Three files, three folders:
ComfyUI/models/diffusion_models/ AniSee-V2-<variant>.safetensors
ComfyUI/models/text_encoders/ qwen_3_06b_base.safetensors
ComfyUI/models/vae/ qwen_image_vae.safetensorsLoad with Load Diffusion Model + Load CLIP + Load VAE — not the Checkpoint Loader. If you already run Anima you have the text encoder and the VAE already. Want the one-file convenience instead? Stay on AIO-V1.
🔬 The Three-Way Comparison
I ran the three V2 variants side by side across ten prompts — seven in natural language, three in pure Danbooru tags — with the seed locked identical across all three models, so the only variable was the checkpoint. Each comparison image in the gallery shows one prompt with all three results next to each other: Base, Turbo, Aesthetic, left to right.
What came out of it:
Turbo diverges the most. On the same seed it lands on a different pose and framing far more often than Base and Aesthetic differ from each other. That is the distillation of the Turbo foundation, not a flaw in the merge — you trade diversity for speed and stability.
Base and Aesthetic are near-twins in composition. Same seed, almost the same layout. The difference is in the finish: Aesthetic has richer colour and softer light, Base is flatter and more neutral.
Tag prompts still work. This one surprised me. The V2 training data is natural language only, with no Danbooru tags at all — yet the base model's tag understanding survives the merge completely intact. Pure tag prompts give you Anima behaviour with an AniSee lean.
Text rendering stays limited. Anima places a single word or a short phrase reasonably well; longer lines come out garbled. Inherited from the base model, and V2 does not improve it.
All ten used sampler
er_sdewith schedulersimpleat 896×1152. Base and Aesthetic ran 40 steps at CFG 4, Turbo 12 steps at CFG 1. Same seed per prompt across all three models.
The ten prompts
Shared negative prompt for all ten:
worst quality, low quality, artist name, blurry, jpeg artifacts,
chromatic aberration, extra fingers, missing fingers, bad hands,
bad anatomy, deformed01 — seesee_elf · Aerial Silk · seed 88123

masterpiece, best quality, safe, seesee_elf, A white-haired elf woman
performs an aerial silk routine in a sunlit forest clearing. She hangs in
a wide split, both hands gripping long emerald fabric strips that fall
from the canopy above. Her platinum blonde hair streams downward, and her
amber eyes look calmly toward the viewer. She wears a form-fitting dark
green bodysuit with fine gold filigree across the chest and thighs, with a
black lace choker at her throat. Warm dappled sunlight filters through the
leaves onto her skin. Full body shot, clean anime rendering with crisp
lines and soft shading.02 — seesee_elf · Library · seed 202

masterpiece, best quality, safe, seesee_elf, A white-haired elf woman
stands in an old library, half turned toward the viewer with a faint
smile. She holds an open leather-bound book in one hand and rests the
other on a tall wooden shelf. Her long platinum hair is loosely braided
over one shoulder and her amber eyes catch the warm lamplight. She wears a
dark green robe with gold embroidery over a cream linen shirt. Dust motes
drift in the beams of afternoon light between towering bookshelves. Upper
body portrait, detailed anime illustration with warm colors.03 — seesee_kitsune · Shrine · seed 303

masterpiece, best quality, safe, seesee_kitsune, An adult fox woman stands
on the stone steps of an autumn shrine, facing the viewer. Tall orange fox
ears with pale inner fur rise above her long straight orange hair, and
several large orange tails with white tips fan out behind her. Her amber
eyes are calm and her expression is gentle. She wears a red and white
shrine maiden kimono with wide sleeves. Red maple leaves fall around her
and a vermilion torii gate stands blurred in the background. Full body,
glossy anime style with warm autumn tones.04 — Pixel Art · Jungle · seed 404

masterpiece, best quality, safe, A single young woman stands in a lush
jungle setting, rendered in pixel art with a visible grid of square
pixels, a limited color palette, and hard-edged blocky shapes defining her
form. She wears a simple explorer outfit with a satchel, and holds a
machete at her side. Broad green fronds and thick vines frame her on both
sides, with a small stone ruin visible behind. Retro 16-bit game aesthetic
with strong outlines and dithered shading.05 — Dragon Woman · Beach · seed 505

masterpiece, best quality, safe, A single adult dragon woman stands on a
black volcanic beach beside a tropical lagoon at sunset, drawn in a clean
modern anime style with crisp cel shading. She rests one hand at her waist
and meets the viewer with a calm, confident expression. Her golden eyes
have narrow pupils, and long wine-red hair falls behind two swept black
horns. Red scales cover her shoulders and outer hips, broad dark crimson
wings rise behind her, and a long plated tail curves around her legs. She
wears a crimson wrap dress with gold trim. Full body, warm orange and deep
blue palette.06 — Painterly · Lighthouse · seed 606

masterpiece, best quality, safe, A digital painting of a lone figure in a
heavy yellow raincoat standing at the base of a stone lighthouse during a
storm. Waves break against the rocks and throw white spray into the air,
and the beam of the lighthouse cuts through low grey cloud. The brushwork
is loose and textured with visible strokes, in a muted palette of slate
blue, ochre and foam white. Wide atmospheric composition, painterly
illustration, not anime.07 — Rooftop Sunset · seed 27182

masterpiece, best quality, safe, An anime girl with short black hair and
green eyes sits on the railing of a city rooftop at sunset, one leg drawn
up and her arms wrapped loosely around her knee. She looks out over a
sprawl of buildings toward an orange and violet sky. She wears a white
school shirt with the sleeves rolled up and a dark pleated skirt, her tie
loosened. Warm rim light catches her hair and shoulders. Cinematic wide
shot with soft gradients and clean line art.08 — Tags · Elf · seed 808

masterpiece, best quality, safe, 1girl, solo, elf, pointy ears, long hair,
white hair, amber eyes, green dress, gold trim, choker, forest, tree,
dappled sunlight, standing, full body, looking at viewer, detailed
background09 — Tags · Fox Girl · seed 909

masterpiece, best quality, safe, 1girl, solo, fox girl, kitsune, animal
ears, fox ears, multiple tails, orange hair, long hair, amber eyes,
japanese clothes, kimono, red skirt, shrine, autumn, maple leaves, smile,
upper body, looking at viewer10 — Tags · Chibi Magical Girl · seed 1010

masterpiece, best quality, safe, 1girl, solo, (chibi:2), magical girl,
pink hair, twintails, blue eyes, frilled dress, star wand, holding wand,
night sky, stars, crescent moon, sparkle, simple background, smile, one
eye closed, full body, wide shot🎛️ Recommended Settings
Sampler er_sde with scheduler simple for every variant. That is my default across the board — neutral style, flat colours, sharp lines.
AniSee-V2-Turbo — 12 steps, CFG 1. The negative prompt has no effect at CFG 1, so don't bother tuning it there.
AniSee-V2-Aesthetic — 40 steps, CFG 4–5. Leave out
score_*tags.AniSee-V2-Base — 40 steps, CFG 4–5.
AniSee AIO-V1 and v1 — 40 steps, CFG 4.5.
CFG guide. 4.0–5.0 is the sweet spot for the 40-step variants. Above 5.0 you start risking burnt images, especially with heavy quality tags. If results feel harsh, drop CFG slightly or cut down on quality tags.
A note on Aesthetic: per the Anima documentation, skip score_* tags entirely on the Aesthetic foundation, in both positive and negative. The base is already high quality and score tags push it into slop territory. masterpiece, best quality is fine to keep.
Sampler alternatives
er_sde+simple— my default. Neutral, flat colours, sharp lines.euler_a— softer, thinner lines, slightly 2.5D, tolerates higher CFG.dpmpp_2m_sde_gpu— similar to er_sde but more creative; can get wild on short prompts.euler— a bit more creative than er_sde. Good on Turbo and Aesthetic, since those are naturally more stable.
📐 Resolution
⭐ Square / general purpose — 1024 × 1024
Portrait / character art — 896 × 1152 or 832 × 1216
Landscape / scenes — 1152 × 896
Wider cinematic — 1254 × 836
Widescreen — 1365 × 768
Stay around 1 MP for the cleanest results. Anima works between 512² and 1536², but starts breaking down somewhere around 2 MP — generate at 1 MP and upscale afterwards if you want bigger.
💡 Prompting
AniSee inherits Anima's prompting system and accepts Danbooru-style tags, natural language, and any mix of the two. A workable structure:
[quality tags] [meta tags] [safety tag] [subject] [character]
[appearance] [pose] [clothing] [background] [lighting] [style]Tag rules inherited from Anima
Lowercase tags, spaces instead of underscores.
Score tags are the only ones using underscores, e.g.
score_7.Artist tags need an
@prefix, e.g.@artistname. Without it the effect is very weak.Where a tag differs between Danbooru and Gelbooru, prefer the Gelbooru spelling.
Prompt weighting works but needs heavier weights than SDXL, e.g.
(chibi:2).
🔑 V2 trigger words
The V2 dataset was captioned in natural language only, so these want to be used in sentences rather than dropped in as bare tags:
seesee_elf— white-haired elfseesee_kitsune— fox woman
Give them a described setting, two sentences minimum. Very short prompts produce unpredictable results on Anima — the model fills the gaps with its own biases, and you may not like what it picks.
✅ Good example — mixed prompt
masterpiece, best quality, score_7, highres, illustration, safe, 1girl,
long silver hair, blue eyes, black hoodie, standing in a rainy city street
at night, neon lights reflecting on wet asphalt, cinematic lighting,
detailed anime illustration✅ Good example — natural language
masterpiece, best quality, score_7, highres, illustration.
A young anime girl with long silver hair and golden eyes, wearing a
traditional shrine maiden outfit with white haori and red hakama.
She stands in a sunlit bamboo forest, cherry blossoms falling softly
around her. Warm afternoon light filtering through the trees,
detailed fabric shading, calm serene expression.❌ Avoid
anime girl, silver hair, hoodieToo sparse. Aim for a decent set of descriptive tags, or two or more sentences.
⭐ Recommended positive prefix
masterpiece, best quality, score_7, highres, illustration,Then your subject, character, scene and style. On Aesthetic, drop the score_7 and use masterpiece, best quality, on its own.
⭐ Recommended negative prompt
worst quality, low quality, score_1, score_2, score_3, artist name,
(lowres:1.2), (worst quality:1.4), (low quality:1.4), (bad anatomy:1.4),
bad hands, multiple views, comic, jpeg artifacts, patreon logo,
patreon username, web address, signature, watermark, artist name,
censored, mosaic censoringIf images come out flat or lose style, ease off the heavy weights — drop (low quality:1.4) back to plain low quality. On Aesthetic, remove the score_* terms. On Turbo at CFG 1 it is ignored entirely.
🛡️ Safety tags
Inherited from Anima. Use one in the positive prompt: safe (recommended default), sensitive, nsfw, or explicit.
🔧 Installation
V2 — any of the three variants
ComfyUI/models/diffusion_models/AniSee-V2-Turbo.safetensors
ComfyUI/models/text_encoders/qwen_3_06b_base.safetensors
ComfyUI/models/vae/qwen_image_vae.safetensorsThen wire up the standard Anima workflow:
Load Diffusion Model →
AniSee-V2-Turbo.safetensorsLoad CLIP →
qwen_3_06b_base.safetensorsLoad VAE →
qwen_image_vae.safetensors
AIO-V1 — one file, one folder
ComfyUI/models/checkpoints/AniSee-AIO-v1.safetensorsLoad it with the standard ComfyUI Checkpoint Loader. Image model, Qwen text encoder and Qwen-Image VAE are all inside — no separate downloads, no missing components, no workflow surgery. At around 5.5 GB it is still smaller than many large anime checkpoints such as Pony, SDXL or ILL-style models, despite packing everything in.
v1 — diffusion model only
ComfyUI/models/diffusion_models/AniSee.safetensors
ComfyUI/models/text_encoders/qwen_3_06b_base.safetensors
ComfyUI/models/vae/qwen_image_vae.safetensorsDrop-in replacement for anima-preview3-base.safetensors in the classic Anima workflow.
📈 Version History
V2 — LoKr merge onto Anima v1.0 / v1.1
Three variants released together: Turbo, Aesthetic, Base
LoKr network — linear 64 / alpha 64, conv 16 / alpha 16, full-rank, factor 4
~24,000 training steps on Anima-Base-v1.0, merged at strength 1.0
Natural-language-only dataset; trigger words
seesee_elfandseesee_kitsuneDiffusion model files only — text encoder and VAE required separately
Moves off Anima Preview3 onto the final Anima v1.0 / v1.1 releases
v1.1 — AniSee AIO
All-in-one checkpoint: image model, Qwen text encoder and Qwen-Image VAE in one file
Loads with the standard ComfyUI Checkpoint Loader
~5.5 GB, goes in
ComfyUI/models/checkpoints/Same AniSee full fine-tune as v1.0
v1.0 — initial release
Full fine-tune of Anima Preview3 Base, ~20,000 steps on a curated anime dataset
LLM adapter only very lightly co-trained, per Anima's fine-tuning guidelines
Diffusion model variant, drop-in for
anima-preview3-base.safetensors
🗺️ Roadmap
✅ Released
AniSee v1 — full fine-tune, diffusion model
AniSee AIO-V1 — all-in-one checkpoint
AniSee V2 — Turbo, Aesthetic and Base
🔜 Planned
AniSee V2 AIO — an all-in-one build of V2, so V2 gets the same one-file convenience AIO-V1 has.
Official AniSee ComfyUI workflow — a dedicated workflow, though the standard Anima workflow already covers both generations.
🖼️ Gallery
A curated gallery of AniSee samples — characters, scenes, different styles and prompts in action. Worth a look before you download.
🎨 View the AniSee Sample Gallery →
🙏 Credits
Base model: Anima by CircleStone Labs and Comfy Org — Preview3 Base for V1, v1.0 / v1.1 for V2
Architecture: built on NVIDIA Cosmos-Predict2-2B; Anima is a derivative model
Fine-tune and merges: SeeSee21
📜 License
AniSee inherits the CircleStone Labs Non-Commercial License from Anima. The model and its derivatives may be used for non-commercial purposes only. As a derivative of Cosmos-Predict2-2B-Text2Image, the NVIDIA Open Model License Agreement also applies insofar as it covers derivative models.
Generated images are not covered by that restriction. You may use the images you make commercially — selling images, paid commissions, concept art or assets for a paid product are all fine. What needs a separate license is hosting the model behind a paid API, embedding the weights in a monetized product, or running it on a paid generation platform.
For commercial licensing of the base model, contact CircleStone Labs at [email protected].
AniSee — a personal anime model built on Anima. 🎨
Description
LoKr merge on anima-base-v1.0 — the same model the LoKr was trained on.
The most neutral and most flexible variant: the widest range of subjects and styles, and the plainest default look. If you want to drive artist tags or push your own style hard, start here.
40 steps, CFG 4–5, sampler er_sde / simple.
This is also the right foundation if you want to train your own LoRAs on AniSee — same architecture as Anima-Base, and CircleStone Labs recommends training LoRAs on the Base variant in every case.
Not an all-in-one file — text encoder and VAE separately.
















