GRAM: Global Reasoning for Multi-Page VQA
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Blau, Tsachi, Fogel, Sharon, Ronen, Roi, Golts, Alona, Ganz, Roy, Avraham, Elad Ben, Aberdam, Aviad, Tsiper, Shahar, Litman, Ron |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DocVLM: Make Your VLM an Efficient Reader
von: Nacson, Mor Shpigel, et al.
Veröffentlicht: (2024)
von: Nacson, Mor Shpigel, et al.
Veröffentlicht: (2024)
Question Aware Vision Transformer for Multimodal Reasoning
von: Ganz, Roy, et al.
Veröffentlicht: (2024)
von: Ganz, Roy, et al.
Veröffentlicht: (2024)
TAP-VL: Text Layout-Aware Pre-training for Enriched Vision-Language Models
von: Fhima, Jonathan, et al.
Veröffentlicht: (2024)
von: Fhima, Jonathan, et al.
Veröffentlicht: (2024)
DODO: Discrete OCR Diffusion Models
von: Man, Sean, et al.
Veröffentlicht: (2026)
von: Man, Sean, et al.
Veröffentlicht: (2026)
Class-Conditioned Transformation for Enhanced Robust Image Classification
von: Blau, Tsachi, et al.
Veröffentlicht: (2023)
von: Blau, Tsachi, et al.
Veröffentlicht: (2023)
VisFocus: Prompt-Guided Vision Encoders for OCR-Free Dense Document Understanding
von: Abramovich, Ofir, et al.
Veröffentlicht: (2024)
von: Abramovich, Ofir, et al.
Veröffentlicht: (2024)
DREAM: Deep Research Evaluation with Agentic Metrics
von: Avraham, Elad Ben, et al.
Veröffentlicht: (2026)
von: Avraham, Elad Ben, et al.
Veröffentlicht: (2026)
Conceptual Learning via Embedding Approximations for Reinforcing Interpretability and Transparency
von: Dikter, Maor, et al.
Veröffentlicht: (2024)
von: Dikter, Maor, et al.
Veröffentlicht: (2024)
Text-to-Image Generation Via Energy-Based CLIP
von: Ganz, Roy, et al.
Veröffentlicht: (2024)
von: Ganz, Roy, et al.
Veröffentlicht: (2024)
Context-aware Prompt Tuning: Advancing In-Context Learning with Adversarial Methods
von: Blau, Tsachi, et al.
Veröffentlicht: (2024)
von: Blau, Tsachi, et al.
Veröffentlicht: (2024)
Enhancing Consistency-Based Image Generation via Adversarialy-Trained Classification and Energy-Based Discrimination
von: Golan, Shelly, et al.
Veröffentlicht: (2024)
von: Golan, Shelly, et al.
Veröffentlicht: (2024)
Learned 3D volumetric recovery of clouds and its uncertainty for climate analysis
von: Ronen, Roi, et al.
Veröffentlicht: (2024)
von: Ronen, Roi, et al.
Veröffentlicht: (2024)
Doc-V*:Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA
von: Zheng, Yuanlei, et al.
Veröffentlicht: (2026)
von: Zheng, Yuanlei, et al.
Veröffentlicht: (2026)
Paint by Inpaint: Learning to Add Image Objects by Removing Them First
von: Wasserman, Navve, et al.
Veröffentlicht: (2024)
von: Wasserman, Navve, et al.
Veröffentlicht: (2024)
AdaptiVocab: Enhancing LLM Efficiency in Focused Domains through Lightweight Vocabulary Adaptation
von: Nakash, Itay, et al.
Veröffentlicht: (2025)
von: Nakash, Itay, et al.
Veröffentlicht: (2025)
DNN-based 3D Cloud Retrieval for Variable Solar Illumination and Multiview Spaceborne Imaging
von: Klein, Tamar, et al.
Veröffentlicht: (2024)
von: Klein, Tamar, et al.
Veröffentlicht: (2024)
DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation
von: Benita, Roi, et al.
Veröffentlicht: (2023)
von: Benita, Roi, et al.
Veröffentlicht: (2023)
Decidability Results for Fragments of First-Order Logic via a Symbolic Model Property
von: Elad, Neta, et al.
Veröffentlicht: (2026)
von: Elad, Neta, et al.
Veröffentlicht: (2026)
Axe 'Em: Eliminating Spurious States with Induction Axioms
von: Elad, Neta, et al.
Veröffentlicht: (2024)
von: Elad, Neta, et al.
Veröffentlicht: (2024)
Pictorial and apictorial polygonal jigsaw puzzles from arbitrary number of crossing cuts
von: Shahar, Peleg Harel Ofir Itzhak, et al.
Veröffentlicht: (2020)
von: Shahar, Peleg Harel Ofir Itzhak, et al.
Veröffentlicht: (2020)
Query complexity lower bounds for local list-decoding and hard-core predicates (even for small rate and huge lists)
von: Ron-Zewi, Noga, et al.
Veröffentlicht: (2024)
von: Ron-Zewi, Noga, et al.
Veröffentlicht: (2024)
FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders
von: Gramaccioni, Riccardo Fosco, et al.
Veröffentlicht: (2025)
von: Gramaccioni, Riccardo Fosco, et al.
Veröffentlicht: (2025)
GRAM-R$^2$: Self-Training Generative Foundation Reward Models for Reward Reasoning
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
von: Natan, Shahar Ben, et al.
Veröffentlicht: (2026)
von: Natan, Shahar Ben, et al.
Veröffentlicht: (2026)
Real-Time 3D Object Detection Using InnovizOne LiDAR and Low-Power Hailo-8 AI Accelerator
von: Krispin-Avraham, Itay, et al.
Veröffentlicht: (2024)
von: Krispin-Avraham, Itay, et al.
Veröffentlicht: (2024)
Solving Convex Partition Visual Jigsaw Puzzles
von: Ohayon, Yaniv, et al.
Veröffentlicht: (2025)
von: Ohayon, Yaniv, et al.
Veröffentlicht: (2025)
PuzLM: Solving Jigsaw Puzzles with Sequence-to-Sequence Language Models
von: Elkin, Gur, et al.
Veröffentlicht: (2025)
von: Elkin, Gur, et al.
Veröffentlicht: (2025)
Pairwise Alignment & Compatibility for Arbitrarily Irregular Image Fragments
von: Shahar, Ofir Itzhak, et al.
Veröffentlicht: (2025)
von: Shahar, Ofir Itzhak, et al.
Veröffentlicht: (2025)
Equilibrium Propagation Without Limits
von: Litman, Elon
Veröffentlicht: (2025)
von: Litman, Elon
Veröffentlicht: (2025)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
von: Shrestha, Robik, et al.
Veröffentlicht: (2020)
von: Shrestha, Robik, et al.
Veröffentlicht: (2020)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
AID-AppEAL: Automatic Image Dataset and Algorithm for Content Appeal Enhancement and Assessment Labeling
von: Chen, Sherry X., et al.
Veröffentlicht: (2024)
von: Chen, Sherry X., et al.
Veröffentlicht: (2024)
Knowledge Condensation and Reasoning for Knowledge-based VQA
von: Hao, Dongze, et al.
Veröffentlicht: (2024)
von: Hao, Dongze, et al.
Veröffentlicht: (2024)
Separating the Wheat from the Chaff: Understanding (In-)Completeness of Proof Mechanisms for Separation Logic with Inductive Definitions
von: Elad, Neta, et al.
Veröffentlicht: (2025)
von: Elad, Neta, et al.
Veröffentlicht: (2025)
An Infinite Needle in a Finite Haystack: Finding Infinite Counter-Models in Deductive Verification
von: Elad, Neta, et al.
Veröffentlicht: (2023)
von: Elad, Neta, et al.
Veröffentlicht: (2023)
The Missing GAP: From Solving Square Jigsaw Puzzles to Handling Real World Archaeological Fragments
von: Shahar, Ofir Itzhak, et al.
Veröffentlicht: (2026)
von: Shahar, Ofir Itzhak, et al.
Veröffentlicht: (2026)
SILO: Solving Inverse Problems with Latent Operators
von: Raphaeli, Ron, et al.
Veröffentlicht: (2025)
von: Raphaeli, Ron, et al.
Veröffentlicht: (2025)
Adversaries With Incentives: A Strategic Alternative to Adversarial Robustness
von: Ehrenberg, Maayan, et al.
Veröffentlicht: (2024)
von: Ehrenberg, Maayan, et al.
Veröffentlicht: (2024)
GRAM: A Generative Foundation Reward Model for Reward Generalization
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
Preprocessing Methods for Memristive Reservoir Computing for Image Recognition
von: Daniels, Rishona, et al.
Veröffentlicht: (2025)
von: Daniels, Rishona, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DocVLM: Make Your VLM an Efficient Reader
von: Nacson, Mor Shpigel, et al.
Veröffentlicht: (2024) -
Question Aware Vision Transformer for Multimodal Reasoning
von: Ganz, Roy, et al.
Veröffentlicht: (2024) -
TAP-VL: Text Layout-Aware Pre-training for Enriched Vision-Language Models
von: Fhima, Jonathan, et al.
Veröffentlicht: (2024) -
DODO: Discrete OCR Diffusion Models
von: Man, Sean, et al.
Veröffentlicht: (2026) -
Class-Conditioned Transformation for Enhanced Robust Image Classification
von: Blau, Tsachi, et al.
Veröffentlicht: (2023)