GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Guan, Yaohan, Wang, Pristina, Dehak, Najim, Yuille, Alan, Chen, Jieneng, Khashabi, Daniel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient Large Multi-modal Models via Visual Context Compression
di: Chen, Jieneng, et al.
Pubblicazione: (2024)
di: Chen, Jieneng, et al.
Pubblicazione: (2024)
Generative World Explorer
di: Lu, Taiming, et al.
Pubblicazione: (2024)
di: Lu, Taiming, et al.
Pubblicazione: (2024)
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
di: Chen, Jieneng, et al.
Pubblicazione: (2024)
di: Chen, Jieneng, et al.
Pubblicazione: (2024)
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
di: Zhang, Tiezheng, et al.
Pubblicazione: (2025)
di: Zhang, Tiezheng, et al.
Pubblicazione: (2025)
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
di: Wang, Xingrui, et al.
Pubblicazione: (2025)
di: Wang, Xingrui, et al.
Pubblicazione: (2025)
Prompt-Based Exemplar Super-Compression and Regeneration for Class-Incremental Learning
di: Duan, Ruxiao, et al.
Pubblicazione: (2023)
di: Duan, Ruxiao, et al.
Pubblicazione: (2023)
SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
di: Ma, Wufei, et al.
Pubblicazione: (2025)
di: Ma, Wufei, et al.
Pubblicazione: (2025)
Computer Vision and Its Relationship to Cognitive Science: A perspective from Bayes Decision Theory
di: Yuille, Alan, et al.
Pubblicazione: (2026)
di: Yuille, Alan, et al.
Pubblicazione: (2026)
GenEx: Generating an Explorable World
di: Lu, Taiming, et al.
Pubblicazione: (2024)
di: Lu, Taiming, et al.
Pubblicazione: (2024)
Thinking with Spatial Code for Physical-World Video Reasoning
di: Chen, Jieneng, et al.
Pubblicazione: (2026)
di: Chen, Jieneng, et al.
Pubblicazione: (2026)
SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference
di: Wang, Feng, et al.
Pubblicazione: (2023)
di: Wang, Feng, et al.
Pubblicazione: (2023)
A Bayesian Approach to OOD Robustness in Image Classification
di: Kaushik, Prakhar, et al.
Pubblicazione: (2024)
di: Kaushik, Prakhar, et al.
Pubblicazione: (2024)
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
di: Wu, Junfei, et al.
Pubblicazione: (2025)
di: Wu, Junfei, et al.
Pubblicazione: (2025)
Cross-Corpus and Cross-domain Handwriting Assessment of NeuroDegenerative Diseases via Time-Series-to-Image Conversion
di: Chavez, Gabrielle, et al.
Pubblicazione: (2025)
di: Chavez, Gabrielle, et al.
Pubblicazione: (2025)
Large Language Models are Universal Reasoners for Visual Generation
di: Ren, Sucheng, et al.
Pubblicazione: (2026)
di: Ren, Sucheng, et al.
Pubblicazione: (2026)
3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
di: Ma, Wufei, et al.
Pubblicazione: (2024)
di: Ma, Wufei, et al.
Pubblicazione: (2024)
Source-Free and Image-Only Unsupervised Domain Adaptation for Category Level Object Pose Estimation
di: Kaushik, Prakhar, et al.
Pubblicazione: (2024)
di: Kaushik, Prakhar, et al.
Pubblicazione: (2024)
Leveraging AI Predicted and Expert Revised Annotations in Interactive Segmentation: Continual Tuning or Full Training?
di: Zhang, Tiezheng, et al.
Pubblicazione: (2024)
di: Zhang, Tiezheng, et al.
Pubblicazione: (2024)
Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
di: Liu, Qihao, et al.
Pubblicazione: (2025)
di: Liu, Qihao, et al.
Pubblicazione: (2025)
PaLM2-VAdapter: Progressively Aligned Language Model Makes a Strong Vision-language Adapter
di: Xiao, Junfei, et al.
Pubblicazione: (2024)
di: Xiao, Junfei, et al.
Pubblicazione: (2024)
CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning
di: Ma, Wenxin, et al.
Pubblicazione: (2026)
di: Ma, Wenxin, et al.
Pubblicazione: (2026)
SPFormer: Enhancing Vision Transformer with Superpixel Representation
di: Mei, Jieru, et al.
Pubblicazione: (2024)
di: Mei, Jieru, et al.
Pubblicazione: (2024)
Are Vision Language Models Ready for Clinical Diagnosis? A 3D Medical Benchmark for Tumor-centric Visual Question Answering
di: Chen, Yixiong, et al.
Pubblicazione: (2025)
di: Chen, Yixiong, et al.
Pubblicazione: (2025)
Detecting Performance Degradation under Data Shift in Pathology Vision-Language Model
di: Guan, Hao, et al.
Pubblicazione: (2026)
di: Guan, Hao, et al.
Pubblicazione: (2026)
Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models
di: He, Jialuo, et al.
Pubblicazione: (2026)
di: He, Jialuo, et al.
Pubblicazione: (2026)
RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba Models
di: Yang, Timing, et al.
Pubblicazione: (2025)
di: Yang, Timing, et al.
Pubblicazione: (2025)
Dynamic Token Reduction during Generation for Vision Language Models
di: Liang, Xiaoyu, et al.
Pubblicazione: (2025)
di: Liang, Xiaoyu, et al.
Pubblicazione: (2025)
QAVA: Query-Agnostic Visual Attack to Large Vision-Language Models
di: Zhang, Yudong, et al.
Pubblicazione: (2025)
di: Zhang, Yudong, et al.
Pubblicazione: (2025)
Skill-Conditioned Visual Geolocation for Vision-Language Models
di: Yang, Chenjie, et al.
Pubblicazione: (2026)
di: Yang, Chenjie, et al.
Pubblicazione: (2026)
Generative Visual Communication in the Era of Vision-Language Models
di: Vinker, Yael
Pubblicazione: (2024)
di: Vinker, Yael
Pubblicazione: (2024)
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
di: Liu, Qihao, et al.
Pubblicazione: (2025)
di: Liu, Qihao, et al.
Pubblicazione: (2025)
4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
di: Zhong, Shanshan, et al.
Pubblicazione: (2025)
di: Zhong, Shanshan, et al.
Pubblicazione: (2025)
World-in-World: World Models in a Closed-Loop World
di: Zhang, Jiahan, et al.
Pubblicazione: (2025)
di: Zhang, Jiahan, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
di: Min, Kyungmin, et al.
Pubblicazione: (2024)
di: Min, Kyungmin, et al.
Pubblicazione: (2024)
Fake it till You Make it: Reward Modeling as Discriminative Prediction
di: Liu, Runtao, et al.
Pubblicazione: (2025)
di: Liu, Runtao, et al.
Pubblicazione: (2025)
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity
di: Xu, Zhenlin, et al.
Pubblicazione: (2023)
di: Xu, Zhenlin, et al.
Pubblicazione: (2023)
Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models
di: Chae, Hyunsik, et al.
Pubblicazione: (2025)
di: Chae, Hyunsik, et al.
Pubblicazione: (2025)
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
di: Wang, Xingrui, et al.
Pubblicazione: (2024)
di: Wang, Xingrui, et al.
Pubblicazione: (2024)
Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models
di: Tan, Huajie, et al.
Pubblicazione: (2025)
di: Tan, Huajie, et al.
Pubblicazione: (2025)
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
di: Wang, Xingrui, et al.
Pubblicazione: (2025)
di: Wang, Xingrui, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Efficient Large Multi-modal Models via Visual Context Compression
di: Chen, Jieneng, et al.
Pubblicazione: (2024) -
Generative World Explorer
di: Lu, Taiming, et al.
Pubblicazione: (2024) -
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
di: Chen, Jieneng, et al.
Pubblicazione: (2024) -
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
di: Zhang, Tiezheng, et al.
Pubblicazione: (2025) -
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
di: Wang, Xingrui, et al.
Pubblicazione: (2025)