CivArchive
    ← All articles
    Published May 28, 2025by AbstractPhila

    Preparing the 1b booru comma style caption only dataset.

    57 views4 reactions1 comments on CivitAI0 collected
    announcement

    Every think to yourself;

    I could use a billion booru captions to train a summarizer specifically devoted to booru stuff - that would otherwise not be possible due to the CAREFUL dataset choices provided by other avenues?


    LOOK NO FURTHER!

    At Phila-Tech we are working to provide all those datasets for your smut-wrangling needs, so you don't have to worry about looking elsewhere. You can simply look to Phila-Tech and it's freely available non-curated absolutely insane datasets, built entirely for the sake of orderly chaos!

    Half of the captions will be plain english, and the other half - ORGANIZED BOORU TAGS!

    The captions will contain realistic and pragmatic combinations of over a billion categorized and SORTED captions; meant to be used to condition and direct your favorite vision model into the direction of your dreams!

    Turn your generic ass plain english, into highly rich and robust implementations of guided insanely wildcarded captions - specifically curated... WITH NO CURATION!

    GOOD LUCK!

    We already have a very fun human single-caption variation - and it's been hard at work training our new shunt conditioniners.

    https://huggingface.co/datasets/AbstractPhil/human-templated-captions-1b

    Not to worry my friends, the Booru version will be ready as soon as work will allow it.

    The booru variation will come equipped with an entire section devoted to teaching your shunts how to tame those misbehaving models as well!

    Don't want to see all that smut? Well, load up the shunt into the negative prompt and away she goes!

    TF is a LECO? Just guide the thing any direction you want and it won't even exist!

    A fantastic and wonderous order created from pure chaos.

    Why... WHY!?

    For pony and illustrious shunts!

    Apparently the plain English versions are falling on deaf ears, however the small trains with booru captions aren't. They are taking like a hot knife through butter.

    So I have a perfect plan; and it includes modifying the baseline response of the clip_l's while these shunts are running, using BERT!

    The Pony and Illustrious versions are going to be four shunts in one - which is a new prototype of high fidelity shunt dispersion using four sets of timesteps.

    One for clip_l positive, one for clip_g positive, one for clip_l negative, and one for clip_g negative.

    FOUR SHUNTS!

    This is gonna be badass. Lets get it.