Saved in:
| Main Authors: | Ataallah, Kirolos, Abdelrahman, Eslam, Ahmed, Mahmoud, Gou, Chenhui, Pahwa, Khushbu, Ding, Jian, Elhoseiny, Mohamed |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2406.19875 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
by: Ataallah, Kirolos, et al.
Published: (2024)
by: Ataallah, Kirolos, et al.
Published: (2024)
Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
by: Ataallah, Kirolos, et al.
Published: (2024)
by: Ataallah, Kirolos, et al.
Published: (2024)
InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity
by: Wang, Haoming, et al.
Published: (2025)
by: Wang, Haoming, et al.
Published: (2025)
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
by: Abdelrahman, Eslam, et al.
Published: (2023)
by: Abdelrahman, Eslam, et al.
Published: (2023)
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
by: Ahmed, Mahmoud, et al.
Published: (2024)
by: Ahmed, Mahmoud, et al.
Published: (2024)
iMotion-LLM: Instruction-Conditioned Trajectory Generation
by: Felemban, Abdulwahab, et al.
Published: (2024)
by: Felemban, Abdulwahab, et al.
Published: (2024)
Analyzing Social Networks of Actors in Movies and TV Shows
by: Giri, Sarthak, et al.
Published: (2024)
by: Giri, Sarthak, et al.
Published: (2024)
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
ToddlerDiffusion: Interactive Structured Image Generation with Cascaded Schrödinger Bridge
by: Abdelrahman, Eslam, et al.
Published: (2023)
by: Abdelrahman, Eslam, et al.
Published: (2023)
3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes
by: Ahmed, Mahmoud, et al.
Published: (2025)
by: Ahmed, Mahmoud, et al.
Published: (2025)
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
by: Barakat, Anas, et al.
Published: (2026)
by: Barakat, Anas, et al.
Published: (2026)
MagnetDB: A Longitudinal Torrent Discovery Dataset with IMDb-Matched Movies and TV Shows
by: Seidenberger, Scott, et al.
Published: (2025)
by: Seidenberger, Scott, et al.
Published: (2025)
How Well Can Vision Language Models See Image Details?
by: Gou, Chenhui, et al.
Published: (2024)
by: Gou, Chenhui, et al.
Published: (2024)
GNNX-BENCH: Unravelling the Utility of Perturbation-based GNN Explainers through In-depth Benchmarking
by: Kosan, Mert, et al.
Published: (2023)
by: Kosan, Mert, et al.
Published: (2023)
FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
by: Khan, Faizan Farooq, et al.
Published: (2025)
by: Khan, Faizan Farooq, et al.
Published: (2025)
Progressive trends in prenatal genetic screening
by: Kirolos Eskandar
Published: (2022)
by: Kirolos Eskandar
Published: (2022)
Liquid biopsy in genitourinary oncology: Current clinical applications and future prospects across prostate, bladder, and renal cancers
by: Kirolos Eskandar
Published: (2025)
by: Kirolos Eskandar
Published: (2025)
Bioimpressão no Transplante de Órgãos: Dos modelos Experimentais às Perspectivas Clínicas
by: Kirolos Eskandar
Published: (2025)
by: Kirolos Eskandar
Published: (2025)
StoryGPT-V: Large Language Models as Consistent Story Visualizers
by: Shen, Xiaoqian, et al.
Published: (2023)
by: Shen, Xiaoqian, et al.
Published: (2023)
MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
by: Wu, Weijia, et al.
Published: (2024)
by: Wu, Weijia, et al.
Published: (2024)
AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?
by: Bao, Han, et al.
Published: (2024)
by: Bao, Han, et al.
Published: (2024)
Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations
by: Haydarov, Kilichbek, et al.
Published: (2023)
by: Haydarov, Kilichbek, et al.
Published: (2023)
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
by: Zheng, Junjie, et al.
Published: (2025)
by: Zheng, Junjie, et al.
Published: (2025)
INSTA-YOLO: Real-Time Instance Segmentation
by: Mohamed, Eslam, et al.
Published: (2021)
by: Mohamed, Eslam, et al.
Published: (2021)
The Morning Show at WLES-TV.
by: Blondell, Beverley
Published: (1979)
by: Blondell, Beverley
Published: (1979)
Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding
by: Zaranis, Emmanouil, et al.
Published: (2025)
by: Zaranis, Emmanouil, et al.
Published: (2025)
Experimental and Computational Fluid Dynamics Analysis of Industrial Water Desalination with High Salinity by Adsorption Chlorine on Resin in a Packed Column
by: Hamideh Mahmoodabadi, et al.
Published: (2024)
by: Hamideh Mahmoodabadi, et al.
Published: (2024)
LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
Multi-modal brain encoding models for multi-modal stimuli
by: Oota, Subba Reddy, et al.
Published: (2025)
by: Oota, Subba Reddy, et al.
Published: (2025)
3DCoMPaT$^{++}$: An improved Large-scale 3D Vision Dataset for Compositional Recognition
by: Slim, Habib, et al.
Published: (2023)
by: Slim, Habib, et al.
Published: (2023)
4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding
by: Zhu, Wenxuan, et al.
Published: (2025)
by: Zhu, Wenxuan, et al.
Published: (2025)
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
by: Kim, Younggun, et al.
Published: (2025)
by: Kim, Younggun, et al.
Published: (2025)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
by: Shen, Xiaoqian, et al.
Published: (2025)
by: Shen, Xiaoqian, et al.
Published: (2025)
Long‐Term Effects of Sugarcane Cultivation on the Physicochemical Quality Indices of Saline and Sodic Soils: A Case Study of South Khuzestan
by: Najm Hashemy, et al.
Published: (2024)
by: Najm Hashemy, et al.
Published: (2024)
Multimodal Chaptering for Long-Form TV Newscast Video
by: Guetari, Khalil, et al.
Published: (2024)
by: Guetari, Khalil, et al.
Published: (2024)
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
by: Shaker, Abdelrahman, et al.
Published: (2025)
by: Shaker, Abdelrahman, et al.
Published: (2025)
Patterns and Dynamics of Netflix TV Show Popularity
by: Lee, Nahyeon, et al.
Published: (2025)
by: Lee, Nahyeon, et al.
Published: (2025)
Evaluation of Bacillus amyloliquefaciensTV‐17C as a potential biocontrol agent for controlling postharvest Penicillium digitatum on orange
by: Meltem Avan, et al.
Published: (2024)
by: Meltem Avan, et al.
Published: (2024)
Video-to-Text Pedestrian Monitoring (VTPM): Leveraging Computer Vision and Large Language Models for Privacy-Preserve Pedestrian Activity Monitoring at Intersections
by: Abdelrahman, Ahmed S., et al.
Published: (2024)
by: Abdelrahman, Ahmed S., et al.
Published: (2024)
KnowMT-Bench: Benchmarking Knowledge-Intensive Long-Form Question Answering in Multi-Turn Dialogues
by: Chen, Junhao, et al.
Published: (2025)
by: Chen, Junhao, et al.
Published: (2025)
Similar Items
-
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
by: Ataallah, Kirolos, et al.
Published: (2024) -
Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
by: Ataallah, Kirolos, et al.
Published: (2024) -
InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity
by: Wang, Haoming, et al.
Published: (2025) -
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
by: Abdelrahman, Eslam, et al.
Published: (2023) -
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
by: Ahmed, Mahmoud, et al.
Published: (2024)