arxiv.org/abs/1803.09219v2
In this paper, a new framework for construction of Cardan grille for information hiding is proposed. Based on the semantic image inpainting technique, the stego image are driven by secret messages directly. A mask called Digital Cardan Grille (DCG) f...
arxiv.org/abs/2401.03407v7
We introduce a novel bilateral reference framework (BiRefNet) for high-resolution dichotomous image segmentation (DIS). It comprises two essential components: the localization module (LM) and the reconstruction module (RM) with our proposed bilateral...
arxiv.org/abs/2602.08700v1
Conversational search systems increasingly employ clarifying questions to refine user queries and improve the search experience. Previous studies have demonstrated the usefulness of text-based clarifying questions in enhancing both retrieval performa...
arxiv.org/abs/2112.01314v3
Image harmonization aims at adjusting the appearance of the foreground to make it more compatible with the background. Without exploring background illumination and its effects on the foreground elements, existing works are incapable of generating a...
arxiv.org/abs/1809.01372v1
Compositing is one of the most important editing operations for images and videos. The process of improving the realism of composite results is often called harmonization. Previous approaches for harmonization mainly focus on images. In this work, we...
arxiv.org/abs/2508.12290v1
The recent growth of large foundation models that can easily generate pseudo-labels for huge quantity of unlabeled data makes unsupervised Zero-Shot Cross-Domain Image Retrieval (UZS-CDIR) less relevant. In this paper, we therefore turn our attention...
arxiv.org/abs/2310.12971v1
The evaluation of machine-generated image captions poses an interesting yet persistent challenge. Effective evaluation measures must consider numerous dimensions of similarity, including semantic relevance, visual structure, object interactions, capt...
arxiv.org/abs/2502.18477v3
Personalization is central to human-AI interaction, yet current diffusion-based image generation systems remain largely insensitive to user diversity. Existing attempts to address this often rely on costly paired preference data or introduce latency...
arxiv.org/abs/2206.05970v3
Adaptive image restoration models can restore images with different degradation levels at inference time without the need to retrain the model. We present an approach that is highly accurate and allows a significant reduction in the number of paramet...
arxiv.org/abs/1804.03312v1
We investigate a novel approach for image restoration by reinforcement learning. Unlike existing studies that mostly train a single large network for a specialized task, we prepare a toolbox consisting of small-scale convolutional networks of differe...
openai.com/index/introducing-4o-image-generation/
Points: 1072 | Comments: 599 | Author: meetpateltech
arxiv.org/abs/2206.01813v1
Most camera images are rendered and saved in the standard RGB (sRGB) format by the camera's hardware. Due to the in-camera photo-finishing routines, nonlinear sRGB images are undesirable for computer vision tasks that assume a direct relationship bet...
arxiv.org/abs/1507.00302v1
We present a method for learning an embedding that places images of humans in similar poses nearby. This embedding can be used as a direct method of comparing images based on human pose, avoiding potential challenges of estimating body joint position...
arxiv.org/abs/2505.22613v1
Image recaptioning is widely used to generate training datasets with enhanced quality for various multimodal tasks. Existing recaptioning methods typically rely on powerful multimodal large language models (MLLMs) to enhance textual descriptions, but...
arxiv.org/abs/2601.00703v2
In digital imaging, image demosaicing is a crucial first step which recovers the RGB information from a color filter array (CFA). Oftentimes, deep learning is utilized to perform image demosaicing. Given that most modern digital imaging applications...
arxiv.org/abs/astro-ph/0402635v1
We report on multi-epoch HST/WFPC2 images of the XZ Tauri binary, and its outflow, covering the period from 1995 to 2001. Data from 1995 to 1998 have already been published in the literature. Additional images, from 1999, 2000 and 2001 are presente...
arxiv.org/abs/2211.04995v1
Background: Increased pericardial adipose tissue (PAT) is associated with many types of cardiovascular disease (CVD). Although cardiac magnetic resonance images (CMRI) are often acquired in patients with CVD, there are currently no tools to automatic...
arxiv.org/abs/2011.11052v1
3D medical image processing with deep learning greatly suffers from a lack of data. Thus, studies carried out in this field are limited compared to works related to 2D natural image analysis, where very large datasets exist. As a result, powerful and...
arxiv.org/abs/2402.16634v1
Skull-stripping is the removal of background and non-brain anatomical features from brain images. While many skull-stripping tools exist, few target pediatric populations. With the emergence of multi-institutional pediatric data acquisition efforts t...
arxiv.org/abs/astro-ph/9804284v1
We consider a general method of deprojecting 2D images to reconstruct the 3D structure of the projected object, assuming axial symmetry. The method consists of the application of the Fourier Slice Theorem to the general case where the axis of symme...