Global Geometry Is Not Enough for Vision Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Chung, Jiwan, Kim, Seon Joo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
APT: Improving Diffusion Models for High Resolution Image Generation with Adaptive Path Tracing
by: Han, Sangmin, et al.
Published: (2025)
by: Han, Sangmin, et al.
Published: (2025)
Object Aware Egocentric Online Action Detection
by: An, Joungbin, et al.
Published: (2024)
by: An, Joungbin, et al.
Published: (2024)
Representing 3D Shapes With 64 Latent Vectors for 3D Diffusion Models
by: Cho, In, et al.
Published: (2025)
by: Cho, In, et al.
Published: (2025)
Frequency-Aware Token Reduction for Efficient Vision Transformer
by: Lee, Dong-Jae, et al.
Published: (2025)
by: Lee, Dong-Jae, et al.
Published: (2025)
Is Geometry Enough? An Evaluation of Landmark-Based Gaze Estimation
by: Agostinelli, Daniele, et al.
Published: (2026)
by: Agostinelli, Daniele, et al.
Published: (2026)
The Geometry of Representational Failures in Vision Language Models
by: Savietto, Daniele, et al.
Published: (2026)
by: Savietto, Daniele, et al.
Published: (2026)
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
by: Kim, Jiwan, et al.
Published: (2025)
by: Kim, Jiwan, et al.
Published: (2025)
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
by: Kim, Junhyeok, et al.
Published: (2025)
by: Kim, Junhyeok, et al.
Published: (2025)
Towards Visual Text Design Transfer Across Languages
by: Choi, Yejin, et al.
Published: (2024)
by: Choi, Yejin, et al.
Published: (2024)
Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!
by: Chung, Jiwan, et al.
Published: (2024)
by: Chung, Jiwan, et al.
Published: (2024)
Is Contrastive Distillation Enough for Learning Comprehensive 3D Representations?
by: Zhang, Yifan, et al.
Published: (2024)
by: Zhang, Yifan, et al.
Published: (2024)
GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation
by: Yang, Jiahao, et al.
Published: (2026)
by: Yang, Jiahao, et al.
Published: (2026)
Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
by: Hyun, Jeongseok, et al.
Published: (2024)
by: Hyun, Jeongseok, et al.
Published: (2024)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
by: Kim, Youngmin, et al.
Published: (2025)
by: Kim, Youngmin, et al.
Published: (2025)
Proximal Vision Transformer: Enhancing Feature Representation through Two-Stage Manifold Geometry
by: Yun, Haoyu, et al.
Published: (2025)
by: Yun, Haoyu, et al.
Published: (2025)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
by: Tian, Kexin, et al.
Published: (2025)
by: Tian, Kexin, et al.
Published: (2025)
Conditional Brownian Bridge Diffusion Model for VHR SAR to Optical Image Translation
by: Kim, Seon-Hoon, et al.
Published: (2024)
by: Kim, Seon-Hoon, et al.
Published: (2024)
SpatialFly: Geometry-Guided Representation Alignment for UAV Vision-and-Language Navigation in Urban Environments
by: Jiang, Wen, et al.
Published: (2026)
by: Jiang, Wen, et al.
Published: (2026)
Patch is Enough: Naturalistic Adversarial Patch against Vision-Language Pre-training Models
by: Kong, Dehong, et al.
Published: (2024)
by: Kong, Dehong, et al.
Published: (2024)
Pygmalion Effect in Vision: Image-to-Clay Translation for Reflective Geometry Reconstruction
by: Lee, Gayoung, et al.
Published: (2025)
by: Lee, Gayoung, et al.
Published: (2025)
HarmoCLIP: Harmonizing Global and Regional Representations in Contrastive Vision-Language Models
by: Zeng, Haoxi, et al.
Published: (2025)
by: Zeng, Haoxi, et al.
Published: (2025)
Dual-Domain Representation Alignment: Bridging 2D and 3D Vision via Geometry-Aware Architecture Search
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts
by: Kim, Hee-Seon, et al.
Published: (2025)
by: Kim, Hee-Seon, et al.
Published: (2025)
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
by: Hyun, Jeongseok, et al.
Published: (2025)
by: Hyun, Jeongseok, et al.
Published: (2025)
Composed Vision-Language Retrieval for Skin Cancer Case Search via Joint Alignment of Global and Local Representations
by: Wang, Yuheng, et al.
Published: (2026)
by: Wang, Yuheng, et al.
Published: (2026)
EvRepSL: Event-Stream Representation via Self-Supervised Learning for Event-Based Vision
by: Qu, Qiang, et al.
Published: (2024)
by: Qu, Qiang, et al.
Published: (2024)
PRISM: Video Dataset Condensation with Progressive Refinement and Insertion for Sparse Motion
by: Choi, Jaehyun, et al.
Published: (2025)
by: Choi, Jaehyun, et al.
Published: (2025)
CarbonNet: How Computer Vision Plays a Role in Climate Change? Application: Learning Geomechanics from Subsurface Geometry of CCS to Mitigate Global Warming
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
Semantic Is Enough: Only Semantic Information For NeRF Reconstruction
by: Wang, Ruibo, et al.
Published: (2024)
by: Wang, Ruibo, et al.
Published: (2024)
Deep Extrinsic Manifold Representation for Vision Tasks
by: Zhang, Tongtong, et al.
Published: (2024)
by: Zhang, Tongtong, et al.
Published: (2024)
Representation Separation for Semantic Segmentation with Vision Transformers
by: Hong, Yuanduo, et al.
Published: (2022)
by: Hong, Yuanduo, et al.
Published: (2022)
RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation
by: Park, Geon, et al.
Published: (2025)
by: Park, Geon, et al.
Published: (2025)
LINR-PCGC: Lossless Implicit Neural Representations for Point Cloud Geometry Compression
by: Huang, Wenjie, et al.
Published: (2025)
by: Huang, Wenjie, et al.
Published: (2025)
GeoWorld-VLM: Geometry from World Models for Vision-Language Models
by: Gu, Renjie, et al.
Published: (2026)
by: Gu, Renjie, et al.
Published: (2026)
GenOL: Generating Diverse Examples for Name-only Online Learning
by: Seo, Minhyuk, et al.
Published: (2024)
by: Seo, Minhyuk, et al.
Published: (2024)
When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs
by: Cao, Fanpu, et al.
Published: (2026)
by: Cao, Fanpu, et al.
Published: (2026)
One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
by: Gao, Yuan, et al.
Published: (2025)
by: Gao, Yuan, et al.
Published: (2025)
Human-Like Coarse Object Representations in Vision Models
by: Gizdov, Andrey, et al.
Published: (2026)
by: Gizdov, Andrey, et al.
Published: (2026)
Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning
by: Bleeker, Maurits, et al.
Published: (2024)
by: Bleeker, Maurits, et al.
Published: (2024)
DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization
by: Lee, Dongyeun, et al.
Published: (2025)
by: Lee, Dongyeun, et al.
Published: (2025)
Similar Items
-
APT: Improving Diffusion Models for High Resolution Image Generation with Adaptive Path Tracing
by: Han, Sangmin, et al.
Published: (2025) -
Object Aware Egocentric Online Action Detection
by: An, Joungbin, et al.
Published: (2024) -
Representing 3D Shapes With 64 Latent Vectors for 3D Diffusion Models
by: Cho, In, et al.
Published: (2025) -
Frequency-Aware Token Reduction for Efficient Vision Transformer
by: Lee, Dong-Jae, et al.
Published: (2025) -
Is Geometry Enough? An Evaluation of Landmark-Based Gaze Estimation
by: Agostinelli, Daniele, et al.
Published: (2026)