arxiv.org/abs/2501.13920v1
With the rapid development of diffusion models, text-to-image(T2I) models have made significant progress, showcasing impressive abilities in prompt following and image generation. Recently launched models such as FLUX.1 and Ideogram2.0, along with ot...
arxiv.org/abs/2110.12442v1
Automatic Image Captioning is the never-ending effort of creating syntactically and validating the accuracy of textual descriptions of an image in natural language with context. The encoder-decoder structure used throughout existing Bengali Image Cap...
arxiv.org/abs/2106.13416v2
Semantic image synthesis is a process for generating photorealistic images from a single semantic mask. To enrich the diversity of multimodal image synthesis, previous methods have controlled the global appearance of an output image by learning a sin...
www.reddit.com/r/StableDiffusion/comments/1pympur/amazing_zimage_workflow_v30_released/
Workflows for **Z-Image-Turbo**, focused on high-quality image styles and user-friendliness. All three workflows have been updated to version 3.0: * [**Amazing Z-Image Workflow**](https://civitai.co...
arxiv.org/abs/1912.13214v1
Image retargeting is a new image processing task that renders the change of aspect ratio in images. One of the most famous image-retargeting algorithms is seam-carving. Although seam-carving is fast and straightforward, it usually distorts the images...
arxiv.org/abs/1806.06357v1
Traditional image steganography often leans interests towards safely embedding hidden information into cover images with payload capacity almost neglected. This paper combines recent deep convolutional neural network methods with image-into-image ste...
github.com/image-rs/image
Encoding and decoding images in Rust (⭐ 5683)
arxiv.org/abs/2511.00046v1
Digital image processing involves the systematic handling of images using advanced computer algorithms, and has gained significant attention in both academic and practical fields. Image enhancement is a crucial preprocessing stage in the image-proces...
arxiv.org/abs/2305.11321v3
Intrinsic Image Decomposition (IID) is a challenging inverse problem that seeks to decompose a natural image into its underlying intrinsic components such as albedo and shading. While recent image decomposition methods rely on learning-based priors o...
arxiv.org/abs/2507.23620v1
Diffusion models have advanced from text-to-image (T2I) to image-to-image (I2I) generation by incorporating structured inputs such as depth maps, enabling fine-grained spatial control. However, existing methods either train separate models for each c...
arxiv.org/abs/2305.13093v2
Recent image restoration methods have produced significant advancements using deep learning. However, existing methods tend to treat the whole image as a single entity, failing to account for the distinct objects in the image that exhibit individual...
arxiv.org/abs/2503.20418v2
This paper introduces ITA-MDT, the Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On (IVTON), designed to overcome the limitations of previous approaches by leveraging the Masked Diffusion Transformer (MDT)...
arxiv.org/abs/1703.10908v4
This paper introduces Quicksilver, a fast deformable image registration method. Quicksilver registration for image-pairs works by patch-wise prediction of a deformation model based directly on image appearance. A deep encoder-decoder network is used...
www.bing.com/ck/a?!&&p=456db665f9df15e0acefe95bb3a3b614b4cca446fc4e2c4d297a136917575a6fJmltdHM9MTc3MjQwOTYwMA&ptn=3&ver=2&hsh=4&fclid=3f0acae5-1e83-6588-2f02-ddf51f1664d4&u=a1aHR0cHM6Ly9lbi5tLndpa2lwZWRpYS5vcmcvd2lraS9XaWtpcGVkaWE6R3JhcGhpY3NfTGFiL1Jlc291cmNlcy9QREZfY29udmVyc2lvbl90b19TVkc&ntb=1
Before learning how to convert PDF images to SVG images it may be useful to learn how to extract images from PDF documents and create PNG, GIF, and JPG images. By using Adobe Reader many …
arxiv.org/abs/2405.12221v3
Spectrograms are 2D representations of sound that look very different from the images found in our visual world. And natural images, when played as spectrograms, make unnatural sounds. In this paper, we show that it is possible to synthesize spectrog...
arxiv.org/abs/2310.19477v2
Recovering clear images from blurry ones with an unknown blur kernel is a challenging problem. Deep image prior (DIP) proposes to use the deep network as a regularizer for a single image rather than as a supervised model, which achieves encouraging r...
arxiv.org/abs/2107.01889v3
Image composition aims to generate realistic composite image by inserting an object from one image into another background image, where the placement (e.g., location, size, occlusion) of inserted object may be unreasonable, which would significantly...
github.com/Enbatamil/Big-Cats-Image-classification
It classify the image whether it is lion/tiger/cheetah/leopard .This image classification can be used in cameras that are fixed in urban and rural areas to identify the roaming of animal and notify to the higher authority to make immediate action to save a lif…
github.com/bob5-tensorslab/skills
TensorLab Skills are AI task definitions for Claude Code, focusing on multimedia generation and processing using TensorsLab's AI models. It provides comprehensive capabilities including text-to-image and image-to-image generation, advanced image editing (avata…
github.com/tensorslab/skills
TensorLab Skills are AI task definitions for Claude Code, focusing on multimedia generation and processing using TensorsLab's AI models. It provides comprehensive capabilities including text-to-image and image-to-image generation, advanced image editing (avata…