Guardado en:
| Autores principales: | Mishra, Abhijit, Li, Mingda, Fu, Hsiang, Noh, Richard, Kim, Minji |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2502.14780 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction
por: Abaskohi, Amirhossein, et al.
Publicado: (2026)
por: Abaskohi, Amirhossein, et al.
Publicado: (2026)
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning
por: Xu, Zhiyang, et al.
Publicado: (2024)
por: Xu, Zhiyang, et al.
Publicado: (2024)
Improved Baselines with Visual Instruction Tuning
por: Liu, Haotian, et al.
Publicado: (2023)
por: Liu, Haotian, et al.
Publicado: (2023)
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
por: Liu, Qihao, et al.
Publicado: (2025)
por: Liu, Qihao, et al.
Publicado: (2025)
InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning
por: Wan, Zifu, et al.
Publicado: (2025)
por: Wan, Zifu, et al.
Publicado: (2025)
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
por: Fan, Zhiwen, et al.
Publicado: (2025)
por: Fan, Zhiwen, et al.
Publicado: (2025)
PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis
por: Lokesh, K, et al.
Publicado: (2026)
por: Lokesh, K, et al.
Publicado: (2026)
Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
por: Kim, Donghoon, et al.
Publicado: (2025)
por: Kim, Donghoon, et al.
Publicado: (2025)
GeoDANO: Geometric VLM with Domain Agnostic Vision Encoder
por: Cho, Seunghyuk, et al.
Publicado: (2025)
por: Cho, Seunghyuk, et al.
Publicado: (2025)
SentinelLMs: Encrypted Input Adaptation and Fine-tuning of Language Models for Private and Secure Inference
por: Mishra, Abhijit, et al.
Publicado: (2023)
por: Mishra, Abhijit, et al.
Publicado: (2023)
Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer
por: Li, Mingda, et al.
Publicado: (2024)
por: Li, Mingda, et al.
Publicado: (2024)
Re:Verse -- Can Your VLM Read a Manga?
por: Baranwal, Aaditya, et al.
Publicado: (2025)
por: Baranwal, Aaditya, et al.
Publicado: (2025)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
por: Liu, Zheng, et al.
Publicado: (2024)
por: Liu, Zheng, et al.
Publicado: (2024)
Unseen from Seen: Rewriting Observation-Instruction Using Foundation Models for Augmenting Vision-Language Navigation
por: Wei, Ziming, et al.
Publicado: (2025)
por: Wei, Ziming, et al.
Publicado: (2025)
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
por: Jiang, Ziyan, et al.
Publicado: (2024)
por: Jiang, Ziyan, et al.
Publicado: (2024)
Infer Induced Sentiment of Comment Response to Video: A New Task, Dataset and Baseline
por: Jia, Qi, et al.
Publicado: (2024)
por: Jia, Qi, et al.
Publicado: (2024)
HPE-CogVLM: Advancing Vision Language Models with a Head Pose Grounding Task
por: Tian, Yu, et al.
Publicado: (2024)
por: Tian, Yu, et al.
Publicado: (2024)
Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration
por: Park, ChaeHun, et al.
Publicado: (2024)
por: Park, ChaeHun, et al.
Publicado: (2024)
Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models
por: Liu, Zikang, et al.
Publicado: (2025)
por: Liu, Zikang, et al.
Publicado: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
por: Du, Yifan, et al.
Publicado: (2023)
por: Du, Yifan, et al.
Publicado: (2023)
VolDoGer: LLM-assisted Datasets for Domain Generalization in Vision-Language Tasks
por: Choi, Juhwan, et al.
Publicado: (2024)
por: Choi, Juhwan, et al.
Publicado: (2024)
VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
por: Meng, Rui, et al.
Publicado: (2025)
por: Meng, Rui, et al.
Publicado: (2025)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
por: Li, Bin, et al.
Publicado: (2025)
por: Li, Bin, et al.
Publicado: (2025)
OViP: Online Vision-Language Preference Learning for VLM Hallucination
por: Liu, Shujun, et al.
Publicado: (2025)
por: Liu, Shujun, et al.
Publicado: (2025)
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
por: Tanaka, Ryota, et al.
Publicado: (2024)
por: Tanaka, Ryota, et al.
Publicado: (2024)
PersonaVLM: Long-Term Personalized Multimodal LLMs
por: Nie, Chang, et al.
Publicado: (2026)
por: Nie, Chang, et al.
Publicado: (2026)
GUIDE: A Guideline-Guided Dataset for Instructional Video Comprehension
por: Liang, Jiafeng, et al.
Publicado: (2024)
por: Liang, Jiafeng, et al.
Publicado: (2024)
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
por: Zhao, Yi, et al.
Publicado: (2026)
por: Zhao, Yi, et al.
Publicado: (2026)
Constructing Multilingual Visual-Text Datasets Revealing Visual Multilingual Ability of Vision Language Models
por: Atuhurra, Jesse, et al.
Publicado: (2024)
por: Atuhurra, Jesse, et al.
Publicado: (2024)
From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
por: Sheta, Hala, et al.
Publicado: (2025)
por: Sheta, Hala, et al.
Publicado: (2025)
Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning
por: Hu, Zhe, et al.
Publicado: (2025)
por: Hu, Zhe, et al.
Publicado: (2025)
LLaVA-OneVision: Easy Visual Task Transfer
por: Li, Bo, et al.
Publicado: (2024)
por: Li, Bo, et al.
Publicado: (2024)
CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
por: Wang, Yanan, et al.
Publicado: (2025)
por: Wang, Yanan, et al.
Publicado: (2025)
Enabling Chatbots with Eyes and Ears: An Immersive Multimodal Conversation System for Dynamic Interactions
por: Jang, Jihyoung, et al.
Publicado: (2025)
por: Jang, Jihyoung, et al.
Publicado: (2025)
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
por: Fu, Xingyu, et al.
Publicado: (2025)
por: Fu, Xingyu, et al.
Publicado: (2025)
SCoPE VLM: Selective Context Processing for Efficient Document Navigation in Vision-Language Models
por: Lim, Gyubeum, et al.
Publicado: (2025)
por: Lim, Gyubeum, et al.
Publicado: (2025)
Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning
por: Singh, Ayush, et al.
Publicado: (2024)
por: Singh, Ayush, et al.
Publicado: (2024)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
por: Zhang, Di, et al.
Publicado: (2024)
por: Zhang, Di, et al.
Publicado: (2024)
Annotation-Free Reinforcement Learning Query Rewriting via Verifiable Search Reward
por: Cha, Sungguk, et al.
Publicado: (2025)
por: Cha, Sungguk, et al.
Publicado: (2025)
IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding
por: Li, Junxian, et al.
Publicado: (2025)
por: Li, Junxian, et al.
Publicado: (2025)
Ejemplares similares
-
ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction
por: Abaskohi, Amirhossein, et al.
Publicado: (2026) -
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning
por: Xu, Zhiyang, et al.
Publicado: (2024) -
Improved Baselines with Visual Instruction Tuning
por: Liu, Haotian, et al.
Publicado: (2023) -
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
por: Liu, Qihao, et al.
Publicado: (2025) -
InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning
por: Wan, Zifu, et al.
Publicado: (2025)