3,157 results for vision · 0.182s

Sponsored Partners
en.wikipedia.org/wiki/Tomori_Kusunoki

Tomori Kusunoki - Wikipedia

"WAR OF THE VISIONS ファイナルファンタジー ブレイブエクスヴィアス 幻影戦争 公式プレイヤーズサイト | SQUARE ENIX". WAR OF THE VISIONS ファイナルファンタジー ブレイブエクスヴィアス 幻影戦争 公式プレイヤーズサイト | SQUARE ENIX

arxiv.org/abs/2512.17394v2

Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?

Theory of Mind (ToM) - the ability to attribute beliefs and intents to others - is fundamental for social intelligence, yet Vision-Language Model (VLM) evaluations remain largely Western-centric. In this work, we introduce CulturalToM-VQA, a benchmar...

arxiv.org/abs/2206.11461v3

Towards Better User Studies in Computer Graphics and Vision

Online crowdsourcing platforms have made it increasingly easy to perform evaluations of algorithm outputs with survey questions like "which image is better, A or B?", leading to their proliferation in vision and graphics research papers. Results of t...

arxiv.org/abs/2103.13023v1

Can Vision Transformers Learn without Natural Images?

Can we complete pre-training of Vision Transformers (ViT) without natural images and human-annotated labels? Although a pre-trained ViT seems to heavily rely on a large-scale dataset and human-annotated labels, recent large-scale datasets contain sev...

github.com/Vision-CAIR/MiniGPT-4

Vision-CAIR/MiniGPT-4

Open-sourced codes for MiniGPT-4 and MiniGPT-v2 (https://minigpt-4.github.io, https://minigpt-v2.github.io/) (⭐ 25760)

arxiv.org/abs/2403.19963v1

Efficient Modulation for Vision Networks

In this work, we present efficient modulation, a novel design for efficient vision networks. We revisit the modulation mechanism, which operates input through convolutional context modeling and feature projection layers, and fuses features via elemen...