arxiv.org/abs/1511.06457v4
Recovering the occlusion relationships between objects is a fundamental human visual ability which yields important information about the 3D world. In this paper we propose a deep network architecture, called DOC, which acts on a single image, detect...
arxiv.org/abs/1104.1326v2
The purpose of this paper is to study the geometry of images of morphisms from Mori dream spaces. First we prove that a variety which admits a surjective morphism from a Mori dream space is again a Mori dream space. Secondly we introduce a natural fa...
arxiv.org/abs/2312.13578v1
The generation of emotional talking faces from a single portrait image remains a significant challenge. The simultaneous achievement of expressive emotional talking and accurate lip-sync is particularly difficult, as expressiveness is often compromis...
arxiv.org/abs/2507.08396v2
Subject-consistent generation (SCG)-aiming to maintain a consistent subject identity across diverse scenes-remains a challenge for text-to-image (T2I) models. Existing training-free SCG methods often achieve consistency at the cost of layout and pose...
github.com/nostra13/Android-Universal-Image-Loader
Powerful and flexible library for loading, caching and displaying images on Android. (⭐ 17066)
arxiv.org/abs/1804.04719v1
Target detection is the front-end stage in any automatic target recognition system for synthetic aperture radar (SAR) imagery (SAR-ATR). The efficacy of the detector directly impacts the succeeding stages in the SAR-ATR processing chain. There are nu...
arxiv.org/abs/1309.5574v1
In this contribution, an image-guided therapy system supporting gynecologic radiation therapy is introduced. The overall workflow of the presented system starts with the arrival of the patient and ends with follow-up examinations by imaging and a sup...
github.com/ayushdabra/dubai-satellite-imagery-segmentation
Multi-Class Semantic Segmentation on Dubai's Satellite Images. (⭐ 87)
github.com/fiji/Stitching
Fiji's Stitching plugins reconstruct big images from tiled input images. (⭐ 114)
arxiv.org/abs/2602.09084v2
We study instruction-based image editing under professional workflows and identify three persistent challenges: (i) editors often over-edit, modifying content beyond the user's intent; (ii) existing models are largely single-turn, while multi-turn ed...
arxiv.org/abs/2303.07679v2
Prediction of image memorability has attracted interest in various fields. Consequently, the prediction accuracy of convolutional neural network (CNN) models has been approaching the empirical upper bound estimated based on human consistency. However...
www.reddit.com/r/MillennialBets/comments/odz6sd/ionq_dmyi_the_leader_in_quantum_computing_15/
**Author**: u/MadeTheAccountForWSB(**Karma:** 4501, **Created:** Mar-2020). [**IonQ ($DMYI): The Leader in Quantum Computing. 15 Reasons and the Bear Case. (Part 1 due to limit on images) on r/spacs...
www.reddit.com/r/MillennialBets/comments/oh0gvk/ionq_dmyi_the_leader_in_quantum_computing_15/
**Author**: u/MadeTheAccountForWSB(**Karma:** 4563, **Created:** Mar-2020). [**IonQ ($DMYI): The Leader in Quantum Computing. 15 Reasons and the Bear Case. (Part 1 due to limit on images) on r/spacs...
arxiv.org/abs/2007.11806v2
Elevator button recognition is a critical function to realize the autonomous operation of elevators. However, challenging image conditions and various image distortions make it difficult to recognize buttons accurately. To fill this gap, we propose a...
arxiv.org/abs/2510.09509v1
We investigate diagonal artifacts present in images captured by several Samsung smartphones and their impact on PRNU-based camera source verification. We first show that certain Galaxy S series models share a common pattern causing fingerprint collis...
arxiv.org/abs/2509.05696v1
Cross-view geo-localization plays a critical role in Unmanned Aerial Vehicle (UAV) localization and navigation. However, significant challenges arise from the drastic viewpoint differences and appearance variations between images. Existing methods pr...
arxiv.org/abs/2112.06482v4
Recently, Multi-modal Named Entity Recognition (MNER) has attracted a lot of attention. Most of the work utilizes image information through region-level visual representations obtained from a pretrained object detector and relies on an attention mech...
arxiv.org/abs/2311.00048v2
Multiple Instance Learning (MIL) has been widely used in weakly supervised whole slide image (WSI) classification. Typical MIL methods include a feature embedding part, which embeds the instances into features via a pre-trained feature extractor, and...
arxiv.org/abs/2208.06049v3
Self-attention based transformer models have been dominating many computer vision tasks in the past few years. Their superb model qualities heavily depend on the excessively large labeled image datasets. In order to reduce the reliance on large label...
github.com/phillipi/pix2pix
Image-to-image translation with conditional adversarial nets (⭐ 10619)