Text-Conditioned Resampler For Long Form Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Korbar, Bruno, Xian, Yongqin, Tonioni, Alessio, Zisserman, Andrew, Tombari, Federico |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
by: Plizzari, Chiara, et al.
Published: (2025)
by: Plizzari, Chiara, et al.
Published: (2025)
Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
by: Korbar, Bruno, et al.
Published: (2025)
by: Korbar, Bruno, et al.
Published: (2025)
LIME: Localized Image Editing via Attention Regularization in Diffusion Models
by: Simsar, Enis, et al.
Published: (2023)
by: Simsar, Enis, et al.
Published: (2023)
UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility Constraint
by: Simsar, Enis, et al.
Published: (2024)
by: Simsar, Enis, et al.
Published: (2024)
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
by: Kuzucu, Selim, et al.
Published: (2026)
by: Kuzucu, Selim, et al.
Published: (2026)
Look, Listen and Recognise: Character-Aware Audio-Visual Subtitling
by: Korbar, Bruno, et al.
Published: (2024)
by: Korbar, Bruno, et al.
Published: (2024)
R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs
by: Xie, Jiahao, et al.
Published: (2026)
by: Xie, Jiahao, et al.
Published: (2026)
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models
by: Xie, Jiahao, et al.
Published: (2026)
by: Xie, Jiahao, et al.
Published: (2026)
Test-Time Visual In-Context Tuning
by: Xie, Jiahao, et al.
Published: (2025)
by: Xie, Jiahao, et al.
Published: (2025)
Active Data Curation Effectively Distills Large-Scale Multimodal Models
by: Udandarao, Vishaal, et al.
Published: (2024)
by: Udandarao, Vishaal, et al.
Published: (2024)
MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
by: Segu, Mattia, et al.
Published: (2025)
by: Segu, Mattia, et al.
Published: (2025)
InseRF: Text-Driven Generative Object Insertion in Neural 3D Scenes
by: Shahbazi, Mohamad, et al.
Published: (2024)
by: Shahbazi, Mohamad, et al.
Published: (2024)
RefAM: Attention Magnets for Zero-Shot Referral Segmentation
by: Kukleva, Anna, et al.
Published: (2025)
by: Kukleva, Anna, et al.
Published: (2025)
Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
by: Bagad, Piyush, et al.
Published: (2025)
by: Bagad, Piyush, et al.
Published: (2025)
Toward a Diffusion-Based Generalist for Dense Vision Tasks
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
BRAVE: Broadening the visual encoding of vision-language models
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
Zero-Shot Styled Text Image Generation, but Make It Autoregressive
by: Pippi, Vittorio, et al.
Published: (2025)
by: Pippi, Vittorio, et al.
Published: (2025)
Training-free Online Video Step Grounding
by: Zanella, Luca, et al.
Published: (2025)
by: Zanella, Luca, et al.
Published: (2025)
Adapting MLLMs for Nuanced Video Retrieval
by: Bagad, Piyush, et al.
Published: (2025)
by: Bagad, Piyush, et al.
Published: (2025)
Autoregressive Styled Text Image Generation, but Make it Reliable
by: Zaccagnino, Carmine, et al.
Published: (2025)
by: Zaccagnino, Carmine, et al.
Published: (2025)
Open-World Object Counting in Videos
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation
by: Bill, Eric Tillmann, et al.
Published: (2026)
by: Bill, Eric Tillmann, et al.
Published: (2026)
Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
by: Go, Hyojun, et al.
Published: (2025)
by: Go, Hyojun, et al.
Published: (2025)
Epipolar Geometry Improves Video Generation Models
by: Kupyn, Orest, et al.
Published: (2025)
by: Kupyn, Orest, et al.
Published: (2025)
Gaussians-to-Life: Text-Driven Animation of 3D Gaussian Splatting Scenes
by: Wimmer, Thomas, et al.
Published: (2024)
by: Wimmer, Thomas, et al.
Published: (2024)
T*: Re-thinking Temporal Search for Long-Form Video Understanding
by: Ye, Jinhui, et al.
Published: (2025)
by: Ye, Jinhui, et al.
Published: (2025)
Zero-Shot Long-Form Video Understanding through Screenplay
by: Wu, Yongliang, et al.
Published: (2024)
by: Wu, Yongliang, et al.
Published: (2024)
Long-VMNet: Accelerating Long-Form Video Understanding via Fixed Memory
by: Gurukar, Saket, et al.
Published: (2025)
by: Gurukar, Saket, et al.
Published: (2025)
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
by: Perrett, Toby, et al.
Published: (2024)
by: Perrett, Toby, et al.
Published: (2024)
The Manga Whisperer: Automatically Generating Transcriptions for Comics
by: Sachdeva, Ragav, et al.
Published: (2024)
by: Sachdeva, Ragav, et al.
Published: (2024)
From Panels to Prose: Generating Literary Narratives from Comics
by: Sachdeva, Ragav, et al.
Published: (2025)
by: Sachdeva, Ragav, et al.
Published: (2025)
A General Protocol to Probe Large Vision Models for 3D Physical Understanding
by: Zhan, Guanqi, et al.
Published: (2023)
by: Zhan, Guanqi, et al.
Published: (2023)
Character-Centric Understanding of Animated Movies
by: Gui, Zhongrui, et al.
Published: (2025)
by: Gui, Zhongrui, et al.
Published: (2025)
Understanding Co-speech Gestures in-the-wild
by: Hegde, Sindhu B, et al.
Published: (2025)
by: Hegde, Sindhu B, et al.
Published: (2025)
HyperSDFusion: Bridging Hierarchical Structures in Language and Geometry for Enhanced 3D Text2Shape Generation
by: Leng, Zhiying, et al.
Published: (2024)
by: Leng, Zhiying, et al.
Published: (2024)
Video Panels for Long Video Understanding
by: Doorenbos, Lars, et al.
Published: (2025)
by: Doorenbos, Lars, et al.
Published: (2025)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
by: Heyward, Joseph, et al.
Published: (2024)
by: Heyward, Joseph, et al.
Published: (2024)
CountGD++: Generalized Prompting for Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
TextVidBench: A Benchmark for Long Video Scene Text Understanding
by: Zhong, Yangyang, et al.
Published: (2025)
by: Zhong, Yangyang, et al.
Published: (2025)
Similar Items
-
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
by: Plizzari, Chiara, et al.
Published: (2025) -
Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
by: Korbar, Bruno, et al.
Published: (2025) -
LIME: Localized Image Editing via Attention Regularization in Diffusion Models
by: Simsar, Enis, et al.
Published: (2023) -
UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility Constraint
by: Simsar, Enis, et al.
Published: (2024) -
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
by: Kuzucu, Selim, et al.
Published: (2026)