From Local Cues to Global Percepts: Emergent Gestalt Organization in Self-Supervised Vision Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Tianqin, Wen, Ziqi, Song, Leiran, Liu, Jun, Jing, Zhi, Lee, Tai Sing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Does resistance to style-transfer equal Global Shape Bias? Measuring network sensitivity to global shape configuration
por: Wen, Ziqi, et al.
Publicado: (2023)
por: Wen, Ziqi, et al.
Publicado: (2023)
Learning More by Seeing Less: Structure First Learning for Efficient, Transferable, and Human-Aligned Vision
por: Li, Tianqin, et al.
Publicado: (2025)
por: Li, Tianqin, et al.
Publicado: (2025)
Perceptual Inductive Bias Is What You Need Before Contrastive Learning
por: Li, Tianqin, et al.
Publicado: (2025)
por: Li, Tianqin, et al.
Publicado: (2025)
DReX: Pure Vision Fusion of Self-Supervised and Convolutional Representations for Image Complexity Prediction
por: Skaza, Jonathan, et al.
Publicado: (2025)
por: Skaza, Jonathan, et al.
Publicado: (2025)
DepthCues: Evaluating Monocular Depth Perception in Large Vision Models
por: Danier, Duolikun, et al.
Publicado: (2024)
por: Danier, Duolikun, et al.
Publicado: (2024)
Modeling Rapid Contextual Learning in the Visual Cortex with Fast-Weight Deep Autoencoder Networks
por: Li, Yue, et al.
Publicado: (2025)
por: Li, Yue, et al.
Publicado: (2025)
Probing and Bridging Geometry-Interaction Cues for Affordance Reasoning in Vision Foundation Models
por: Zhang, Qing, et al.
Publicado: (2026)
por: Zhang, Qing, et al.
Publicado: (2026)
Weakly-Supervised Image Forgery Localization via Vision-Language Collaborative Reasoning Framework
por: Sheng, Ziqi, et al.
Publicado: (2025)
por: Sheng, Ziqi, et al.
Publicado: (2025)
Leveraging Semantic Cues from Foundation Vision Models for Enhanced Local Feature Correspondence
por: Cadar, Felipe, et al.
Publicado: (2024)
por: Cadar, Felipe, et al.
Publicado: (2024)
CIEC: Coupling Implicit and Explicit Cues for Multimodal Weakly Supervised Manipulation Localization
por: Yu, Xinquan, et al.
Publicado: (2026)
por: Yu, Xinquan, et al.
Publicado: (2026)
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
por: Li, Yifan, et al.
Publicado: (2025)
por: Li, Yifan, et al.
Publicado: (2025)
Self-Attention-Based Contextual Modulation Improves Neural System Identification
por: Lin, Isaac, et al.
Publicado: (2024)
por: Lin, Isaac, et al.
Publicado: (2024)
Exploiting Polarized Material Cues for Robust Car Detection
por: Dong, Wen, et al.
Publicado: (2024)
por: Dong, Wen, et al.
Publicado: (2024)
Visual Cues of Gender and Race are Associated with Stereotyping in Vision-Language Models
por: Lee, Messi H. J., et al.
Publicado: (2025)
por: Lee, Messi H. J., et al.
Publicado: (2025)
PatchCue: Enhancing Vision-Language Model Reasoning with Patch-Based Visual Cues
por: Qi, Yukun, et al.
Publicado: (2026)
por: Qi, Yukun, et al.
Publicado: (2026)
Self-Supervised Weight Templates for Scalable Vision Model Initialization
por: Xie, Yucheng, et al.
Publicado: (2026)
por: Xie, Yucheng, et al.
Publicado: (2026)
Emergent Bayesian Behaviour and Optimal Cue Combination in LLMs
por: Ma, Julian, et al.
Publicado: (2025)
por: Ma, Julian, et al.
Publicado: (2025)
From Perception to Action: An Interactive Benchmark for Vision Reasoning
por: Wu, Yuhao, et al.
Publicado: (2026)
por: Wu, Yuhao, et al.
Publicado: (2026)
Generalization of Self-Supervised Vision Transformers for Protein Localization Across Microscopy Domains
por: Isselmann, Ben, et al.
Publicado: (2026)
por: Isselmann, Ben, et al.
Publicado: (2026)
PP-SSL : Priority-Perception Self-Supervised Learning for Fine-Grained Recognition
por: Li, ShuaiHeng, et al.
Publicado: (2024)
por: Li, ShuaiHeng, et al.
Publicado: (2024)
LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization
por: Huang, Zhenpeng, et al.
Publicado: (2026)
por: Huang, Zhenpeng, et al.
Publicado: (2026)
Self-Localized Collaborative Perception
por: Ni, Zhenyang, et al.
Publicado: (2024)
por: Ni, Zhenyang, et al.
Publicado: (2024)
Self-Supervised Vision Transformers Are Efficient Segmentation Learners for Imperfect Labels
por: Lee, Seungho, et al.
Publicado: (2024)
por: Lee, Seungho, et al.
Publicado: (2024)
Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation
por: Lu, Junxin, et al.
Publicado: (2026)
por: Lu, Junxin, et al.
Publicado: (2026)
From Local to Global to Mechanistic: An iERF-Centered Unified Framework for Interpreting Vision Models
por: Kim, Yearim, et al.
Publicado: (2026)
por: Kim, Yearim, et al.
Publicado: (2026)
From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection
por: Cai, Lincan, et al.
Publicado: (2025)
por: Cai, Lincan, et al.
Publicado: (2025)
All-in-one Weather-degraded Image Restoration via Adaptive Degradation-aware Self-prompting Model
por: Wen, Yuanbo, et al.
Publicado: (2024)
por: Wen, Yuanbo, et al.
Publicado: (2024)
GLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video Segmentation
por: Lin, Lang, et al.
Publicado: (2025)
por: Lin, Lang, et al.
Publicado: (2025)
Local-to-Global Self-Supervised Representation Learning for Diabetic Retinopathy Grading
por: Hajighasemlou, Mostafa, et al.
Publicado: (2024)
por: Hajighasemlou, Mostafa, et al.
Publicado: (2024)
Efficient Self-supervised Vision Pretraining with Local Masked Reconstruction
por: Chen, Jun, et al.
Publicado: (2022)
por: Chen, Jun, et al.
Publicado: (2022)
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
por: Feng, X., et al.
Publicado: (2024)
por: Feng, X., et al.
Publicado: (2024)
Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation
por: Yang, Longzhen, et al.
Publicado: (2025)
por: Yang, Longzhen, et al.
Publicado: (2025)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
por: Lu, Yifan, et al.
Publicado: (2025)
por: Lu, Yifan, et al.
Publicado: (2025)
Hierarchical Self-Supervised Adversarial Training for Robust Vision Models in Histopathology
por: Malik, Hashmat Shadab, et al.
Publicado: (2025)
por: Malik, Hashmat Shadab, et al.
Publicado: (2025)
AtlasVid: Efficient Ultra-High-Resolution Long Video Generation via Decoupled Global-Local Modeling
por: Mai, Ziyang, et al.
Publicado: (2026)
por: Mai, Ziyang, et al.
Publicado: (2026)
Stealthy Backdoor Attack in Self-Supervised Learning Vision Encoders for Large Vision Language Models
por: Liu, Zhaoyi, et al.
Publicado: (2025)
por: Liu, Zhaoyi, et al.
Publicado: (2025)
LoDisc: Learning Global-Local Discriminative Features for Self-Supervised Fine-Grained Visual Recognition
por: Shi, Jialu, et al.
Publicado: (2024)
por: Shi, Jialu, et al.
Publicado: (2024)
Self-Supervised Sparse Sensor Fusion for Long Range Perception
por: Palladin, Edoardo, et al.
Publicado: (2025)
por: Palladin, Edoardo, et al.
Publicado: (2025)
Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking
por: Wen, Wen, et al.
Publicado: (2025)
por: Wen, Wen, et al.
Publicado: (2025)
Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding
por: Murlidaran, Shravan, et al.
Publicado: (2026)
por: Murlidaran, Shravan, et al.
Publicado: (2026)
Ejemplares similares
-
Does resistance to style-transfer equal Global Shape Bias? Measuring network sensitivity to global shape configuration
por: Wen, Ziqi, et al.
Publicado: (2023) -
Learning More by Seeing Less: Structure First Learning for Efficient, Transferable, and Human-Aligned Vision
por: Li, Tianqin, et al.
Publicado: (2025) -
Perceptual Inductive Bias Is What You Need Before Contrastive Learning
por: Li, Tianqin, et al.
Publicado: (2025) -
DReX: Pure Vision Fusion of Self-Supervised and Convolutional Representations for Image Complexity Prediction
por: Skaza, Jonathan, et al.
Publicado: (2025) -
DepthCues: Evaluating Monocular Depth Perception in Large Vision Models
por: Danier, Duolikun, et al.
Publicado: (2024)