Why Relational Graphs Will Save the Next Generation of Vision Foundation Models?
Fuente:
arXiv
Guardado en:
| Autor principal: | Ziaeetabar, Fatemeh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Leveraging Foundation Models for Multimodal Graph-Based Action Recognition
por: Ziaeetabar, Fatemeh, et al.
Publicado: (2025)
por: Ziaeetabar, Fatemeh, et al.
Publicado: (2025)
EfficientGFormer: Multimodal Brain Tumor Segmentation via Pruned Graph-Augmented Transformer
por: Ziaeetabar, Fatemeh
Publicado: (2025)
por: Ziaeetabar, Fatemeh
Publicado: (2025)
Neuro-Symbolic Manipulation Understanding with Enriched Semantic Event Chains
por: Ziaeetabar, Fatemeh
Publicado: (2026)
por: Ziaeetabar, Fatemeh
Publicado: (2026)
Beyond Sequences: A Benchmark for Atomic Hand-Object Interaction Using a Static RNN Encoder
por: Movahed, Yousef Azizi, et al.
Publicado: (2025)
por: Movahed, Yousef Azizi, et al.
Publicado: (2025)
DRL-Guided Neural Batch Sampling for Semi-Supervised Pixel-Level Anomaly Detection
por: Noghredeh, Amirhossein Khadivi, et al.
Publicado: (2025)
por: Noghredeh, Amirhossein Khadivi, et al.
Publicado: (2025)
Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis
por: Qu, Jingguo, et al.
Publicado: (2025)
por: Qu, Jingguo, et al.
Publicado: (2025)
RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence
por: He, Xuming, et al.
Publicado: (2025)
por: He, Xuming, et al.
Publicado: (2025)
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
por: Guo, Jianyuan, et al.
Publicado: (2024)
por: Guo, Jianyuan, et al.
Publicado: (2024)
AD-SAM: Fine-Tuning the Segment Anything Vision Foundation Model for Autonomous Driving Perception
por: Camarena, Mario, et al.
Publicado: (2025)
por: Camarena, Mario, et al.
Publicado: (2025)
Vision Foundation Models as Generalist Tokenizers for Image Generation
por: Zheng, Anlin, et al.
Publicado: (2026)
por: Zheng, Anlin, et al.
Publicado: (2026)
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
por: Wang, Chenting, et al.
Publicado: (2025)
por: Wang, Chenting, et al.
Publicado: (2025)
Pix2Next: Leveraging Vision Foundation Models for RGB to NIR Image Translation
por: Jin, Youngwan, et al.
Publicado: (2024)
por: Jin, Youngwan, et al.
Publicado: (2024)
FSOD-VFM: Few-Shot Object Detection with Vision Foundation Models and Graph Diffusion
por: Feng, Chen-Bin, et al.
Publicado: (2026)
por: Feng, Chen-Bin, et al.
Publicado: (2026)
M-SpecGene: Generalized Foundation Model for RGBT Multispectral Vision
por: Zhou, Kailai, et al.
Publicado: (2025)
por: Zhou, Kailai, et al.
Publicado: (2025)
Rein++: Efficient Generalization and Adaptation for Semantic Segmentation with Vision Foundation Models
por: Wei, Zhixiang, et al.
Publicado: (2025)
por: Wei, Zhixiang, et al.
Publicado: (2025)
Set Pivot Learning: Redefining Generalized Segmentation with Vision Foundation Models
por: Li, Xinhui, et al.
Publicado: (2025)
por: Li, Xinhui, et al.
Publicado: (2025)
Are Vision Foundation Models Foundational for Electron Microscopy Image Segmentation?
por: Fuster-Barceló, Caterina, et al.
Publicado: (2026)
por: Fuster-Barceló, Caterina, et al.
Publicado: (2026)
Sapiens: Foundation for Human Vision Models
por: Khirodkar, Rawal, et al.
Publicado: (2024)
por: Khirodkar, Rawal, et al.
Publicado: (2024)
Mamba as a Bridge: Where Vision Foundation Models Meet Vision Language Models for Domain-Generalized Semantic Segmentation
por: Zhang, Xin, et al.
Publicado: (2025)
por: Zhang, Xin, et al.
Publicado: (2025)
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
por: An, Xiang, et al.
Publicado: (2026)
por: An, Xiang, et al.
Publicado: (2026)
Implicit Modeling for Transferability Estimation of Vision Foundation Models
por: Zheng, Yaoyan, et al.
Publicado: (2025)
por: Zheng, Yaoyan, et al.
Publicado: (2025)
Explainability for Vision Foundation Models: A Survey
por: Kazmierczak, Rémi, et al.
Publicado: (2025)
por: Kazmierczak, Rémi, et al.
Publicado: (2025)
Low-Resource Vision Challenges for Foundation Models
por: Zhang, Yunhua, et al.
Publicado: (2024)
por: Zhang, Yunhua, et al.
Publicado: (2024)
Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic Segmentation
por: Wei, Zhixiang, et al.
Publicado: (2023)
por: Wei, Zhixiang, et al.
Publicado: (2023)
Understanding and Improving Training-Free AI-Generated Image Detections with Vision Foundation Models
por: Tsai, Chung-Ting, et al.
Publicado: (2024)
por: Tsai, Chung-Ting, et al.
Publicado: (2024)
From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models
por: Li, Rongjie, et al.
Publicado: (2024)
por: Li, Rongjie, et al.
Publicado: (2024)
VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
por: Zhang, Xiangdong, et al.
Publicado: (2025)
por: Zhang, Xiangdong, et al.
Publicado: (2025)
GraphVL: Graph-Enhanced Semantic Modeling via Vision-Language Models for Generalized Class Discovery
por: Solanki, Bhupendra, et al.
Publicado: (2024)
por: Solanki, Bhupendra, et al.
Publicado: (2024)
Cracks in the Foundation: A Civil Infrastructure Dataset to Challenge Vision Foundation Models
por: Farronato, Nicola, et al.
Publicado: (2026)
por: Farronato, Nicola, et al.
Publicado: (2026)
EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next
por: Pan, Ye, et al.
Publicado: (2026)
por: Pan, Ye, et al.
Publicado: (2026)
Style-Pro: Style-Guided Prompt Learning for Generalizable Vision-Language Models
por: Talemi, Niloufar Alipour, et al.
Publicado: (2024)
por: Talemi, Niloufar Alipour, et al.
Publicado: (2024)
SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models
por: Makarov, Vladislav, et al.
Publicado: (2026)
por: Makarov, Vladislav, et al.
Publicado: (2026)
FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized Segmentation
por: Zhao, Dong, et al.
Publicado: (2025)
por: Zhao, Dong, et al.
Publicado: (2025)
ED-SAM: An Efficient Diffusion Sampling Approach to Domain Generalization in Vision-Language Foundation Models
por: Truong, Thanh-Dat, et al.
Publicado: (2024)
por: Truong, Thanh-Dat, et al.
Publicado: (2024)
TAP into the Patch Tokens: Leveraging Vision Foundation Model Features for AI-Generated Image Detection
por: Abdullah, Ahmed, et al.
Publicado: (2026)
por: Abdullah, Ahmed, et al.
Publicado: (2026)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
por: Chen, Zhe, et al.
Publicado: (2023)
por: Chen, Zhe, et al.
Publicado: (2023)
CyCLeGen: Cycle-Consistent Layout Prediction and Image Generation in Vision Foundation Models
por: Shan, Xiaojun, et al.
Publicado: (2026)
por: Shan, Xiaojun, et al.
Publicado: (2026)
CanViT: Toward Active-Vision Foundation Models
por: Berreby, Yohaï-Eliel, et al.
Publicado: (2026)
por: Berreby, Yohaï-Eliel, et al.
Publicado: (2026)
Bootstrapping SparseFormers from Vision Foundation Models
por: Gao, Ziteng, et al.
Publicado: (2023)
por: Gao, Ziteng, et al.
Publicado: (2023)
Annotation Free Semantic Segmentation with Vision Foundation Models
por: Seifi, Soroush, et al.
Publicado: (2024)
por: Seifi, Soroush, et al.
Publicado: (2024)
Ejemplares similares
-
Leveraging Foundation Models for Multimodal Graph-Based Action Recognition
por: Ziaeetabar, Fatemeh, et al.
Publicado: (2025) -
EfficientGFormer: Multimodal Brain Tumor Segmentation via Pruned Graph-Augmented Transformer
por: Ziaeetabar, Fatemeh
Publicado: (2025) -
Neuro-Symbolic Manipulation Understanding with Enriched Semantic Event Chains
por: Ziaeetabar, Fatemeh
Publicado: (2026) -
Beyond Sequences: A Benchmark for Atomic Hand-Object Interaction Using a Static RNN Encoder
por: Movahed, Yousef Azizi, et al.
Publicado: (2025) -
DRL-Guided Neural Batch Sampling for Semi-Supervised Pixel-Level Anomaly Detection
por: Noghredeh, Amirhossein Khadivi, et al.
Publicado: (2025)