Re:Verse -- Can Your VLM Read a Manga?
Fuente:
arXiv
Guardado en:
| Autores principales: | Baranwal, Aaditya, Kataria, Madhav, Agrawal, Naitik, Rawat, Yogesh S, Vyas, Shruti |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MolSight: Molecular Property Prediction with Images
por: Baranwal, Aaditya, et al.
Publicado: (2026)
por: Baranwal, Aaditya, et al.
Publicado: (2026)
MolVision: Molecular Property Prediction with Vision Language Models
por: Adak, Deepan, et al.
Publicado: (2025)
por: Adak, Deepan, et al.
Publicado: (2025)
iSafetyBench: A video-language benchmark for safety in industrial environment
por: Abdullah, Raiyaan, et al.
Publicado: (2025)
por: Abdullah, Raiyaan, et al.
Publicado: (2025)
Advancing Automatic Photovoltaic Defect Detection using Semi-Supervised Semantic Segmentation of Electroluminescence Images
por: Jha, Abhishek, et al.
Publicado: (2024)
por: Jha, Abhishek, et al.
Publicado: (2024)
LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models
por: Pathak, Priyank, et al.
Publicado: (2025)
por: Pathak, Priyank, et al.
Publicado: (2025)
SynSpill: Improved Industrial Spill Detection With Synthetic Data
por: Baranwal, Aaditya, et al.
Publicado: (2025)
por: Baranwal, Aaditya, et al.
Publicado: (2025)
ReVersion: Diffusion-Based Relation Inversion from Images
por: Huang, Ziqi, et al.
Publicado: (2023)
por: Huang, Ziqi, et al.
Publicado: (2023)
Semi-supervised Active Learning for Video Action Detection
por: Singh, Ayush, et al.
Publicado: (2023)
por: Singh, Ayush, et al.
Publicado: (2023)
ChemPro: A Progressive Chemistry Benchmark for Large Language Models
por: Baranwal, Aaditya, et al.
Publicado: (2026)
por: Baranwal, Aaditya, et al.
Publicado: (2026)
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
por: Baek, Jeonghun, et al.
Publicado: (2025)
por: Baek, Jeonghun, et al.
Publicado: (2025)
Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding
por: Baek, Jeonghun, et al.
Publicado: (2026)
por: Baek, Jeonghun, et al.
Publicado: (2026)
Colors See Colors Ignore: Clothes Changing ReID with Color Disentanglement
por: Pathak, Priyank, et al.
Publicado: (2025)
por: Pathak, Priyank, et al.
Publicado: (2025)
DIFFER: Disentangling Identity Features via Semantic Cues for Clothes-Changing Person Re-ID
por: Liang, Xin, et al.
Publicado: (2025)
por: Liang, Xin, et al.
Publicado: (2025)
Coarse Attribute Prediction with Task Agnostic Distillation for Real World Clothes Changing ReID
por: Pathak, Priyank, et al.
Publicado: (2025)
por: Pathak, Priyank, et al.
Publicado: (2025)
BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs
por: Baranwal, Aaditya, et al.
Publicado: (2026)
por: Baranwal, Aaditya, et al.
Publicado: (2026)
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
por: Li, Chaoyu, et al.
Publicado: (2025)
por: Li, Chaoyu, et al.
Publicado: (2025)
Bridging Foundation Models and ASTM Metallurgical Standards for Automated Grain Size Estimation from Microscopy Images
por: Mueez, Abdul, et al.
Publicado: (2026)
por: Mueez, Abdul, et al.
Publicado: (2026)
MapVerse: A Benchmark for Geospatial Question Answering on Diverse Real-World Maps
por: Bhat, Sharat, et al.
Publicado: (2026)
por: Bhat, Sharat, et al.
Publicado: (2026)
Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning
por: Singh, Ayush, et al.
Publicado: (2024)
por: Singh, Ayush, et al.
Publicado: (2024)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
por: Zhang, Di, et al.
Publicado: (2024)
por: Zhang, Di, et al.
Publicado: (2024)
DisenQ: Disentangling Q-Former for Activity-Biometrics
por: Azad, Shehreen, et al.
Publicado: (2025)
por: Azad, Shehreen, et al.
Publicado: (2025)
ProDiG: Progressive Diffusion-Guided Gaussian Splatting for Aerial to Ground Reconstruction
por: Mitra, Sirshapan, et al.
Publicado: (2026)
por: Mitra, Sirshapan, et al.
Publicado: (2026)
VIBE: Can a VLM Read the Room?
por: Chakraborty, Tania, et al.
Publicado: (2025)
por: Chakraborty, Tania, et al.
Publicado: (2025)
Understanding Museum Exhibits using Vision-Language Reasoning
por: Balauca, Ada-Astrid, et al.
Publicado: (2024)
por: Balauca, Ada-Astrid, et al.
Publicado: (2024)
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
por: Zhang, Renrui, et al.
Publicado: (2024)
por: Zhang, Renrui, et al.
Publicado: (2024)
Activity-Biometrics: Person Identification from Daily Activities
por: Azad, Shehreen, et al.
Publicado: (2024)
por: Azad, Shehreen, et al.
Publicado: (2024)
VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
por: Kumar, Divake, et al.
Publicado: (2026)
por: Kumar, Divake, et al.
Publicado: (2026)
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
por: Zhao, Yi, et al.
Publicado: (2026)
por: Zhao, Yi, et al.
Publicado: (2026)
Automated LaTeX Code Generation from Handwritten Math Expressions Using Vision Transformer
por: Sundararaj, Jayaprakash, et al.
Publicado: (2024)
por: Sundararaj, Jayaprakash, et al.
Publicado: (2024)
ReVision: A Dataset and Baseline VLM for Privacy-Preserving Task-Oriented Visual Instruction Rewriting
por: Mishra, Abhijit, et al.
Publicado: (2025)
por: Mishra, Abhijit, et al.
Publicado: (2025)
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
por: Kang, Caixin, et al.
Publicado: (2025)
por: Kang, Caixin, et al.
Publicado: (2025)
PersonaVLM: Long-Term Personalized Multimodal LLMs
por: Nie, Chang, et al.
Publicado: (2026)
por: Nie, Chang, et al.
Publicado: (2026)
Systematic Reward Gap Optimization for Mitigating VLM Hallucinations
por: He, Lehan, et al.
Publicado: (2024)
por: He, Lehan, et al.
Publicado: (2024)
GeoDANO: Geometric VLM with Domain Agnostic Vision Encoder
por: Cho, Seunghyuk, et al.
Publicado: (2025)
por: Cho, Seunghyuk, et al.
Publicado: (2025)
EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents
por: Li, Xinze, et al.
Publicado: (2026)
por: Li, Xinze, et al.
Publicado: (2026)
OViP: Online Vision-Language Preference Learning for VLM Hallucination
por: Liu, Shujun, et al.
Publicado: (2025)
por: Liu, Shujun, et al.
Publicado: (2025)
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
por: Siingh, Shikhhar, et al.
Publicado: (2025)
por: Siingh, Shikhhar, et al.
Publicado: (2025)
STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding
por: Garg, Aaryan, et al.
Publicado: (2025)
por: Garg, Aaryan, et al.
Publicado: (2025)
Scaling Open-Vocabulary Action Detection
por: Sia, Zhen Hao, et al.
Publicado: (2025)
por: Sia, Zhen Hao, et al.
Publicado: (2025)
Navigating Hallucinations for Reasoning of Unintentional Activities
por: Grover, Shresth, et al.
Publicado: (2024)
por: Grover, Shresth, et al.
Publicado: (2024)
Ejemplares similares
-
MolSight: Molecular Property Prediction with Images
por: Baranwal, Aaditya, et al.
Publicado: (2026) -
MolVision: Molecular Property Prediction with Vision Language Models
por: Adak, Deepan, et al.
Publicado: (2025) -
iSafetyBench: A video-language benchmark for safety in industrial environment
por: Abdullah, Raiyaan, et al.
Publicado: (2025) -
Advancing Automatic Photovoltaic Defect Detection using Semi-Supervised Semantic Segmentation of Electroluminescence Images
por: Jha, Abhishek, et al.
Publicado: (2024) -
LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models
por: Pathak, Priyank, et al.
Publicado: (2025)