Saved in:
| Main Authors: | Haydarov, Kilichbek, Shen, Xiaoqian, Madasu, Avinash, Salem, Mahmoud, Li, Li-Jia, Elsayed, Gamaleldin, Elhoseiny, Mohamed |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2308.16349 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StoryGPT-V: Large Language Models as Consistent Story Visualizers
by: Shen, Xiaoqian, et al.
Published: (2023)
by: Shen, Xiaoqian, et al.
Published: (2023)
A Shared Valence Axis Across Modern LLMs and Human EEG: The Saturation Regularity
by: Radwan, Yousef A., et al.
Published: (2026)
by: Radwan, Yousef A., et al.
Published: (2026)
No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages
by: Mohamed, Youssef, et al.
Published: (2024)
by: Mohamed, Youssef, et al.
Published: (2024)
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
by: Abdelrahman, Eslam, et al.
Published: (2023)
by: Abdelrahman, Eslam, et al.
Published: (2023)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
by: Shen, Xiaoqian, et al.
Published: (2025)
by: Shen, Xiaoqian, et al.
Published: (2025)
3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes
by: Ahmed, Mahmoud, et al.
Published: (2025)
by: Ahmed, Mahmoud, et al.
Published: (2025)
ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models
by: Madasu, Avinash, et al.
Published: (2023)
by: Madasu, Avinash, et al.
Published: (2023)
VISAFF: Speaker-Centered Visual Affective Feature Learning for Emotion Recognition in Conversation
by: ZHU, Linan, et al.
Published: (2026)
by: ZHU, Linan, et al.
Published: (2026)
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
by: Ataallah, Kirolos, et al.
Published: (2024)
by: Ataallah, Kirolos, et al.
Published: (2024)
Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time
by: Chowdhury, Sanjoy, et al.
Published: (2024)
by: Chowdhury, Sanjoy, et al.
Published: (2024)
Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling
by: Ye, Zilyu, et al.
Published: (2024)
by: Ye, Zilyu, et al.
Published: (2024)
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
by: Ahmed, Mahmoud, et al.
Published: (2024)
by: Ahmed, Mahmoud, et al.
Published: (2024)
Cultural Awareness in Vision-Language Models: A Cross-Country Exploration
by: Madasu, Avinash, et al.
Published: (2025)
by: Madasu, Avinash, et al.
Published: (2025)
Pruning the Paradox: How CLIP's Most Informative Heads Enhance Performance While Amplifying Bias
by: Madasu, Avinash, et al.
Published: (2025)
by: Madasu, Avinash, et al.
Published: (2025)
Affective Flow Language Model for Emotional Support Conversation
by: Zou, Chenghui, et al.
Published: (2026)
by: Zou, Chenghui, et al.
Published: (2026)
Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review
by: Yu, Sungduk, et al.
Published: (2025)
by: Yu, Sungduk, et al.
Published: (2025)
Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
by: Yuan, Haobo, et al.
Published: (2025)
by: Yuan, Haobo, et al.
Published: (2025)
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation
by: Li, Linfei, et al.
Published: (2026)
by: Li, Linfei, et al.
Published: (2026)
Affective Visualization Design: Leveraging the Emotional Impact of Data
by: Lan, Xingyu, et al.
Published: (2023)
by: Lan, Xingyu, et al.
Published: (2023)
Aurelia: Test-time Reasoning Distillation in Audio-Visual LLMs
by: Chowdhury, Sanjoy, et al.
Published: (2025)
by: Chowdhury, Sanjoy, et al.
Published: (2025)
Learning from Reasoning Failures via Synthetic Data Generation
by: Stan, Gabriela Ben Melech, et al.
Published: (2025)
by: Stan, Gabriela Ben Melech, et al.
Published: (2025)
InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
by: Ataallah, Kirolos, et al.
Published: (2024)
by: Ataallah, Kirolos, et al.
Published: (2024)
Quantifying and Enabling the Interpretability of CLIP-like Models
by: Madasu, Avinash, et al.
Published: (2024)
by: Madasu, Avinash, et al.
Published: (2024)
iMotion-LLM: Instruction-Conditioned Trajectory Generation
by: Felemban, Abdulwahab, et al.
Published: (2024)
by: Felemban, Abdulwahab, et al.
Published: (2024)
AgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding
by: Li, Haocheng, et al.
Published: (2026)
by: Li, Haocheng, et al.
Published: (2026)
RVTBench: A Benchmark for Visual Reasoning Tasks
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
by: Chowdhury, Sanjoy, et al.
Published: (2025)
by: Chowdhury, Sanjoy, et al.
Published: (2025)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
by: Shen, Xiaoqian, et al.
Published: (2025)
by: Shen, Xiaoqian, et al.
Published: (2025)
Affective Color Scales for Colormap Data Visualizations
by: Braun, Halle C., et al.
Published: (2025)
by: Braun, Halle C., et al.
Published: (2025)
Neurotoxicidad en neonatos con hiperbilirrubinemia severa. Análisis de los factores de riesgo para neurotoxicidad en neonatos con ictericia severa
by: Rasha Gamaleldin
Published: (2012)
by: Rasha Gamaleldin
Published: (2012)
KOREYS TILIDA O'LCHOV BIRLIKLARI NOMLARINING TARKIBIY TUZILISHI VA HOSIL BO'LISH XUSUSIYATLARI
by: Haydarov, Jasur
Published: (2025)
by: Haydarov, Jasur
Published: (2025)
BOSHQARUV TIZIMINI TAKOMILLASHTIRISHDA AXBOROT-KOMMUNIKATSION TEXNOLOGIYALARNING O'RNI
by: Husanboy Haydarov
Published: (2025)
by: Husanboy Haydarov
Published: (2025)
Seeing Culture: A Benchmark for Visual Reasoning and Grounding
by: Satar, Burak, et al.
Published: (2025)
by: Satar, Burak, et al.
Published: (2025)
AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
by: Park, Se Jin, et al.
Published: (2024)
by: Park, Se Jin, et al.
Published: (2024)
Beyond Static Visual Tokens: Structured Sequential Visual Chain-of-Thought Reasoning
by: Guo, Guangfu, et al.
Published: (2026)
by: Guo, Guangfu, et al.
Published: (2026)
Enhancing Visual Dialog State Tracking through Iterative Object-Entity Alignment in Multi-Round Conversations
by: Pang, Wei, et al.
Published: (2024)
by: Pang, Wei, et al.
Published: (2024)
VGR: Visual Grounded Reasoning
by: Wang, Jiacong, et al.
Published: (2025)
by: Wang, Jiacong, et al.
Published: (2025)
Adaptive Masking Enhances Visual Grounding
by: Jia, Sen, et al.
Published: (2024)
by: Jia, Sen, et al.
Published: (2024)
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
by: Lu, Lidong, et al.
Published: (2025)
by: Lu, Lidong, et al.
Published: (2025)
Similar Items
-
StoryGPT-V: Large Language Models as Consistent Story Visualizers
by: Shen, Xiaoqian, et al.
Published: (2023) -
A Shared Valence Axis Across Modern LLMs and Human EEG: The Saturation Regularity
by: Radwan, Yousef A., et al.
Published: (2026) -
No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages
by: Mohamed, Youssef, et al.
Published: (2024) -
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
by: Abdelrahman, Eslam, et al.
Published: (2023) -
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
by: Shen, Xiaoqian, et al.
Published: (2025)