arxiv.org/abs/1911.09943v2
Recent studies have shown how disentangling images into content and feature spaces can provide controllable image translation/ manipulation. In this paper, we propose a framework to enable utilizing discrete multi-labels to control which features to...
arxiv.org/abs/2309.13601v2
Automated image caption generation is essential for improving the accessibility and understanding of visual content. In this study, we introduce FaceGemma, a model that accurately describes facial attributes such as emotions, expressions, and feature...
arxiv.org/abs/2008.04200v1
Manipulating visual attributes of images through human-written text is a very challenging task. On the one hand, models have to learn the manipulation without the ground truth of the desired output. On the other hand, models have to deal with the inh...
www.reddit.com/r/inIndiannews/comments/1qj14fp/apple_maps_have_updated_their_satellite_imagery/
...
arxiv.org/abs/2501.06031v2
Transductive zero-shot learning with vision-language models leverages image-image similarities within the dataset to achieve better classification accuracy compared to the inductive setting. However, there is little work that explores the structure o...
arxiv.org/abs/2303.12394v1
Road extraction is a process of automatically generating road maps mainly from satellite images. Existing models all target to generate roads from the scratch despite that a large quantity of road maps, though incomplete, are publicly available (e.g....
arxiv.org/abs/2512.01498v2
This report presents solutions to three machine learning challenges developed as part of the Rayan AI Contest: compositional image retrieval, zero-shot anomaly detection, and backdoored model detection. In compositional image retrieval, we developed...
arxiv.org/abs/2601.22125v2
Creative image generation has emerged as a compelling area of research, driven by the need to produce novel and high-quality images that expand the boundaries of imagination. In this work, we propose a novel framework for creative generation using di...
arxiv.org/abs/2104.11931v1
We propose an approach to generate images of people given a desired appearance and pose. Disentangled representations of pose and appearance are necessary to handle the compound variability in the resulting generated images. Hence, we develop an appr...
arxiv.org/abs/astro-ph/9503096v1
Gravitational light deflection can distort the images of distant sources by its tidal effects. The population of faint blue galaxies is at sufficiently high redshift so that their images are distorted near foreground clusters, with giant luminous a...
arxiv.org/abs/2602.21877v1
Image memorability, i.e., how likely an image is to be remembered, has traditionally been studied in computer vision either as a passive prediction task, with models regressing a scalar score, or with generative methods altering the visual input to b...
stackoverflow.com/questions/49669561/animating-images-using-json-and-xml
Tags: android, json, xml, animation, video-processing | Score: -1
arxiv.org/abs/2108.03541v2
Images tell powerful stories but cannot always be trusted. Matching images back to trusted sources (attribution) enables users to make a more informed judgment of the images they encounter online. We propose a robust image hashing algorithm to perfor...
arxiv.org/abs/2004.06165v5
Large-scale pre-training methods of learning cross-modal representations on image-text pairs are becoming popular for vision-language tasks. While existing methods simply concatenate image region features and text features as input to the model to be...
arxiv.org/abs/2602.21712v1
Tooth image segmentation is a cornerstone of dental digitization. However, traditional image encoders relying on fixed-resolution feature maps often lead to discontinuous segmentation and poor discrimination between target regions and background, due...
github.com/leejet/stable-diffusion.cpp
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++ (⭐ 5525)
arxiv.org/abs/2304.10244v2
While lightweight ViT framework has made tremendous progress in image super-resolution, its uni-dimensional self-attention modeling, as well as homogeneous aggregation scheme, limit its effective receptive field (ERF) to include more comprehensive in...
www.reddit.com/r/StableDiffusion/comments/1odsid9/2000s_analog_core_a_hi8_camcorder_lora_for/
Hey, everyone ? I’m excited to share my new LoRA (this time for Qwen-Image), **2000s Analog Core**. I've put a ton of effort and passion into this model. It's designed to perfectly replicat...
en.wikipedia.org/wiki/Public_image_of_John_McCain
Senator John McCain's personal character has dominated the image and perception of him. His family's military heritage, his rebellious nature as a youth
en.wikipedia.org/wiki/Computer-generated_imagery
Computer-generated imagery (CGI) is a specific application of computer graphics for creating or improving images in art, printed media, simulators, videos