12,132 results for Image

arxiv.org/abs/1511.06457v4

DOC: Deep OCclusion Estimation From a Single Image

Recovering the occlusion relationships between objects is a fundamental human visual ability which yields important information about the 3D world. In this paper we propose a deep network architecture, called DOC, which acts on a single image, detect...

arxiv.org/abs/1104.1326v2

On images of Mori dream spaces

The purpose of this paper is to study the geometry of images of morphisms from Mori dream spaces. First we prove that a variety which admits a surjective morphism from a Mori dream space is again a Mori dream space. Secondly we introduce a natural fa...

arxiv.org/abs/2507.08396v2

CoDi: Subject-Consistent and Pose-Diverse Text-to-Image Generation

Subject-consistent generation (SCG)-aiming to maintain a consistent subject identity across diverse scenes-remains a challenge for text-to-image (T2I) models. Existing training-free SCG methods often achieve consistency at the cost of layout and pose...

github.com/fiji/Stitching

fiji/Stitching

Fiji's Stitching plugins reconstruct big images from tiled input images. (⭐ 114)

arxiv.org/abs/2303.07679v2

Feature representations useful for predicting image memorability

Prediction of image memorability has attracted interest in various fields. Consequently, the prediction accuracy of convolutional neural network (CNN) models has been approaching the empirical upper bound estimated based on human consistency. However...

arxiv.org/abs/2510.09509v1

Diagonal Artifacts in Samsung Images: PRNU Challenges and Solutions

We investigate diagonal artifacts present in images captured by several Samsung smartphones and their impact on PRNU-based camera source verification. We first show that certain Galaxy S series models share a common pattern causing fingerprint collis...

arxiv.org/abs/2112.06482v4

ITA: Image-Text Alignments for Multi-Modal Named Entity Recognition

Recently, Multi-modal Named Entity Recognition (MNER) has attracted a lot of attention. Most of the work utilizes image information through region-level visual representations obtained from a pretrained object detector and relies on an attention mech...

arxiv.org/abs/2208.06049v3

MILAN: Masked Image Pretraining on Language Assisted Representation

Self-attention based transformer models have been dominating many computer vision tasks in the past few years. Their superb model qualities heavily depend on the excessively large labeled image datasets. In order to reduce the reliance on large label...

github.com/phillipi/pix2pix

phillipi/pix2pix

Image-to-image translation with conditional adversarial nets (⭐ 10619)