arxiv.org/abs/2506.13496v3
Patent images are technical drawings that convey information about a patent's innovation. Patent image retrieval systems aim to search in vast collections and retrieve the most relevant images. Despite recent advances in information retrieval, patent...
arxiv.org/abs/2411.17237v5
The development of multimodal large language models (MLLMs) enables the evaluation of image quality through natural language descriptions. This advancement allows for more detailed assessments. However, these MLLM-based IQA methods primarily rely on...
github.com/singlestore-labs/singlestoredb-dev-image
The SingleStoreDB Dev Container is the fastest way to develop with SingleStore on your laptop or in a CI/CD environment. (⭐ 62)
github.com/imagentleman/chindogu.php
A quick and dirty one file high performance MVC PHP micro framework (⭐ 5)
arxiv.org/abs/2508.20751v1
Recent advancements highlight the importance of GRPO-based reinforcement learning methods and benchmarking in enhancing text-to-image (T2I) generation. However, current methods using pointwise reward models (RM) for scoring generated images are susce...
arxiv.org/abs/0909.2017v5
A property of sparse representations in relation to their capacity for information storage is discussed. It is shown that this feature can be used for an application that we term Encrypted Image Folding. The proposed procedure is realizable through a...
arxiv.org/abs/2004.08052v2
In this paper, we have trained several deep convolutional networks with introduced training techniques for classifying X-ray images into three classes: normal, pneumonia, and COVID-19, based on two open-source datasets. Our data contains 180 X-ray im...
arxiv.org/abs/1808.05205v1
Deep learning has shown promising results in medical image analysis, however, the lack of very large annotated datasets confines its full potential. Although transfer learning with ImageNet pre-trained classification models can alleviate the problem,...
arxiv.org/abs/2203.09457v1
Novel view synthesis from a single image has recently attracted a lot of attention, and it has been primarily advanced by 3D deep learning and rendering techniques. However, most work is still limited by synthesizing new views within relatively small...
arxiv.org/abs/1004.0766v1
Separation of the text regions from background texture and graphics is an important step of any optical character recognition sytem for the images containg both texts and graphics. In this paper, we have presented a novel text/graphics separation tec...
arxiv.org/abs/2310.02642v1
Event cameras are a type of novel neuromorphic sen-sor that has been gaining increasing attention. Existing event-based backbones mainly rely on image-based designs to extract spatial information within the image transformed from events, overlooking...
arxiv.org/abs/2312.12659v1
Recent advances in vision language pretraining (VLP) have been largely attributed to the large-scale data collected from the web. However, uncurated dataset contains weakly correlated image-text pairs, causing data inefficiency. To address the issue,...
arxiv.org/abs/2103.13023v1
Can we complete pre-training of Vision Transformers (ViT) without natural images and human-annotated labels? Although a pre-trained ViT seems to heavily rely on a large-scale dataset and human-annotated labels, recent large-scale datasets contain sev...
www.reddit.com/r/Atari2600/comments/1rjj0kx/repair_2600_distorted_image/
Recently bought a 2600 4switch, seller told me it was functional. Bought two games from another seller but sadly the image is showing vertical lines. I can start the games, I don't know if the gamepla...
arxiv.org/abs/2004.03264v3
As machine learning for images becomes democratized in the Software 2.0 era, one of the serious bottlenecks is securing enough labeled data for training. This problem is especially critical in a manufacturing setting where smart factories rely on mac...
www.reddit.com/r/CharacterAI/comments/1mn300m/this_image_violated_the_image_guidelines/
...
arxiv.org/abs/2207.04873v2
Image Retrieval is commonly evaluated with Average Precision (AP) or Recall@k. Yet, those metrics, are limited to binary labels and do not take into account errors' severity. This paper introduces a new hierarchical AP training method for pertinent i...
arxiv.org/abs/1910.09233v1
Peoples nowadays prefer to use digital gadgets like cameras or mobile phones for capturing documents. Automatic extraction of panels/characters from the images of a comic document is challenging due to the wide variety of drawing styles adopted by wr...
www.bing.com/ck/a?!&&p=e72aa16ce590b51384aaef186140ad801e36bee427c8c662b3fc0654d517b366JmltdHM9MTc3MjQ5NjAwMA&ptn=3&ver=2&hsh=4&fclid=10d2ade0-044d-6bd8-016b-baf105216abb&u=a1aHR0cHM6Ly93d3cuaXN0b2NrcGhvdG8uY29tL3Bob3Rvcy9mZWV0P21zb2NraWQ9MTBkMmFkZTAwNDRkNmJkODAxNmJiYWYxMDUyMTZhYmI&ntb=1
Search from 1,447,253 Feet stock photos, pictures and royalty-free images from iStock. For the first time, get 1 free month of iStock exclusive photos, illustrations, and more.
arxiv.org/abs/2503.01294v1
In this paper, we propose a novel garment-centric outpainting (GCO) framework based on the latent diffusion model (LDM) for fine-grained controllable apparel showcase image generation. The proposed framework aims at customizing a fashion model wearin...