Showing posts with label LAION-400M model. Show all posts
Showing posts with label LAION-400M model. Show all posts

Thursday, July 14, 2022

north shore deep image prior

 



LAION CLIP model inversion image synthesis. So in theory you are seeing a visualization of what the CLIP model is representing for a particular 'maui north shore scene' text prompt.

Thursday, April 28, 2022

the Brutality of War

 

Blow up in Studio Artist with paint action sequence processing.  From an animation, but blogger is balking at the video up load for some reason.

Babbler

 

One interesting behavior with the LAION-400M latent diffusion generative ai model is that it generates very weird images when you give it a gibberish text prompt.  This one was prompted with 'arghhh'.  Who are these people anyway and why do they come up when you feed the model this text?

Other nonsense prompts give you something like this where you get random people along with a mangled version of the gibberish text prompt built into the generated image.

Changing the gibberish prompt slightly can push you into a different space where you get mangled text mixed with the weird alien cartoony characters that are very characteristic of this particular model.

Using 'zappp' gives you mangled text built into some weird abstract design structure.  Probing the boundaries of the system in this way exposes to some extent what the synthesis algorithm is doing.  It helps in the analysis of this to think of generative texture models, how they are built, what is going on under the hood, and what the resulting output looks like.  It also gives you all kinds of clues to building alternate algorithms for synthesis that are not neural net based but would give similar looking results.


 

Tuesday, April 26, 2022

the End of Time - a children's story

 


A few grabs from a generative ai guided drift session using latent diffusion with the LION-400M public model.

Complete guided generative diffusion session animated in Studio Artist.


Monday, April 25, 2022

maui hana highway traffic pileup

 

This one is hilarious if you live here.  It also tells you oh so much about what is going on in the synthesis algorithm.

Also makes me wonder about the landscape renditions they love so much in the paper.  The stitching works for this kind of thing.

Does not work so much for this 512 x 512 rendition of Hotel Street in Chinatown in Honolulu.
The 256 x 256 output seems to be much more coherent, but once you go above that it reminds me of what you see going on in VQGAN.

Same thing with this shot at rendering Honolulu harbor.  The VQGAN model i've been working with does a much better job at this kind of thing.  The LAION-400M dataset and associated CLIP seems to not be as good for Hawaii specific localisms.  Both VQGAN and the VQVAE model i tried seemed to be better at catching local style idiosyncrasies i like to mess with.

The other thing about this model is that a hallucinated stock photo watermark footer seems to be rendered by the generative algorithm at the bottom of the generated image approx every 10 images.  They should reallu clean that stuff out of the database before training on it (i think).