Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ataallah, Kirolos, Shen, Xiaoqian, Abdelrahman, Eslam, Sleiman, Essam, Zhuge, Mingchen, Ding, Jian, Zhu, Deyao, Schmidhuber, Jürgen, Elhoseiny, Mohamed |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024)
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024)
InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024)
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024)
iMotion-LLM: Instruction-Conditioned Trajectory Generation
von: Felemban, Abdulwahab, et al.
Veröffentlicht: (2024)
von: Felemban, Abdulwahab, et al.
Veröffentlicht: (2024)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2025)
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2025)
StoryGPT-V: Large Language Models as Consistent Story Visualizers
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2023)
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2023)
Beyond Outlining: Heterogeneous Recursive Planning for Adaptive Long-form Writing with Language Models
von: Xiong, Ruibin, et al.
Veröffentlicht: (2025)
von: Xiong, Ruibin, et al.
Veröffentlicht: (2025)
Small Vision-Language Models are Smart Compressors for Long Video Understanding
von: Fei, Junjie, et al.
Veröffentlicht: (2026)
von: Fei, Junjie, et al.
Veröffentlicht: (2026)
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
von: Abdelrahman, Eslam, et al.
Veröffentlicht: (2023)
von: Abdelrahman, Eslam, et al.
Veröffentlicht: (2023)
Language Agents as Optimizable Graphs
von: Zhuge, Mingchen, et al.
Veröffentlicht: (2024)
von: Zhuge, Mingchen, et al.
Veröffentlicht: (2024)
How Well Can Vision Language Models See Image Details?
von: Gou, Chenhui, et al.
Veröffentlicht: (2024)
von: Gou, Chenhui, et al.
Veröffentlicht: (2024)
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2024)
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2024)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2025)
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2025)
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2024)
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2024)
ToddlerDiffusion: Interactive Structured Image Generation with Cascaded Schrödinger Bridge
von: Abdelrahman, Eslam, et al.
Veröffentlicht: (2023)
von: Abdelrahman, Eslam, et al.
Veröffentlicht: (2023)
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
von: Kim, Younggun, et al.
Veröffentlicht: (2025)
von: Kim, Younggun, et al.
Veröffentlicht: (2025)
Time Blindness: Why Video-Language Models Can't See What Humans Can?
von: Upadhyay, Ujjwal, et al.
Veröffentlicht: (2025)
von: Upadhyay, Ujjwal, et al.
Veröffentlicht: (2025)
MiniGPT-Med: Large Language Model as a General Interface for Radiology Diagnosis
von: Alkhaldi, Asma, et al.
Veröffentlicht: (2024)
von: Alkhaldi, Asma, et al.
Veröffentlicht: (2024)
Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
von: Wang, Wenyi, et al.
Veröffentlicht: (2025)
von: Wang, Wenyi, et al.
Veröffentlicht: (2025)
ScholarChemQA: Unveiling the Power of Language Models in Chemical Research Question Answering
von: Chen, Xiuying, et al.
Veröffentlicht: (2024)
von: Chen, Xiuying, et al.
Veröffentlicht: (2024)
Liquid biopsy in genitourinary oncology: Current clinical applications and future prospects across prostate, bladder, and renal cancers
von: Kirolos Eskandar
Veröffentlicht: (2025)
von: Kirolos Eskandar
Veröffentlicht: (2025)
Bioimpressão no Transplante de Órgãos: Dos modelos Experimentais às Perspectivas Clínicas
von: Kirolos Eskandar
Veröffentlicht: (2025)
von: Kirolos Eskandar
Veröffentlicht: (2025)
Progressive trends in prenatal genetic screening
von: Kirolos Eskandar
Veröffentlicht: (2022)
von: Kirolos Eskandar
Veröffentlicht: (2022)
FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
von: Khan, Faizan Farooq, et al.
Veröffentlicht: (2025)
von: Khan, Faizan Farooq, et al.
Veröffentlicht: (2025)
3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2025)
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2025)
Goldfish: Monolingual Language Models for 350 Languages
von: Chang, Tyler A., et al.
Veröffentlicht: (2024)
von: Chang, Tyler A., et al.
Veröffentlicht: (2024)
INSTA-YOLO: Real-Time Instance Segmentation
von: Mohamed, Eslam, et al.
Veröffentlicht: (2021)
von: Mohamed, Eslam, et al.
Veröffentlicht: (2021)
Video-to-Text Pedestrian Monitoring (VTPM): Leveraging Computer Vision and Large Language Models for Privacy-Preserve Pedestrian Activity Monitoring at Intersections
von: Abdelrahman, Ahmed S., et al.
Veröffentlicht: (2024)
von: Abdelrahman, Ahmed S., et al.
Veröffentlicht: (2024)
Fast Camouflaged Object Detection via Edge-based Reversible Re-calibration Network
von: Ji, Ge-Peng, et al.
Veröffentlicht: (2021)
von: Ji, Ge-Peng, et al.
Veröffentlicht: (2021)
Goldfish: An Efficient Federated Unlearning Framework
von: Wang, Houzhe, et al.
Veröffentlicht: (2024)
von: Wang, Houzhe, et al.
Veröffentlicht: (2024)
Document Haystacks: Vision-Language Reasoning Over Piles of 1000+ Documents
von: Chen, Jun, et al.
Veröffentlicht: (2024)
von: Chen, Jun, et al.
Veröffentlicht: (2024)
Experimental and Computational Fluid Dynamics Analysis of Industrial Water Desalination with High Salinity by Adsorption Chlorine on Resin in a Packed Column
von: Hamideh Mahmoodabadi, et al.
Veröffentlicht: (2024)
von: Hamideh Mahmoodabadi, et al.
Veröffentlicht: (2024)
EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture
von: Gamil, Mohamed, et al.
Veröffentlicht: (2025)
von: Gamil, Mohamed, et al.
Veröffentlicht: (2025)
Efficient Self-supervised Vision Pretraining with Local Masked Reconstruction
von: Chen, Jun, et al.
Veröffentlicht: (2022)
von: Chen, Jun, et al.
Veröffentlicht: (2022)
MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models
von: Ahmed, Seif, et al.
Veröffentlicht: (2025)
von: Ahmed, Seif, et al.
Veröffentlicht: (2025)
Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations
von: Haydarov, Kilichbek, et al.
Veröffentlicht: (2023)
von: Haydarov, Kilichbek, et al.
Veröffentlicht: (2023)
TapToTab : Video-Based Guitar Tabs Generation using AI and Audio Analysis
von: Ghaleb, Ali, et al.
Veröffentlicht: (2024)
von: Ghaleb, Ali, et al.
Veröffentlicht: (2024)
dTRPO: Trajectory Reduction in Policy Optimization of Diffusion Large Language Models
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2026)
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2026)
Long‐Term Effects of Sugarcane Cultivation on the Physicochemical Quality Indices of Saline and Sodic Soils: A Case Study of South Khuzestan
von: Najm Hashemy, et al.
Veröffentlicht: (2024)
von: Najm Hashemy, et al.
Veröffentlicht: (2024)
RPRA: Predicting an LLM-Judge for Efficient but Performant Inference
von: Ashley, Dylan R., et al.
Veröffentlicht: (2026)
von: Ashley, Dylan R., et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024) -
InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024) -
iMotion-LLM: Instruction-Conditioned Trajectory Generation
von: Felemban, Abdulwahab, et al.
Veröffentlicht: (2024) -
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2025) -
StoryGPT-V: Large Language Models as Consistent Story Visualizers
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2023)