Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data
Fuente:
arXiv
Salvato in:
| Autori principali: | Ma, Wufei, Li, Kai, Jiang, Zhongshi, Meshry, Moustafa, Liu, Qihao, Wang, Huiyu, Häne, Christian, Yuille, Alan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ImageNet3D: Towards General-Purpose Object-Level 3D Understanding
di: Ma, Wufei, et al.
Pubblicazione: (2024)
di: Ma, Wufei, et al.
Pubblicazione: (2024)
TriDiff-4D: Fast 4D Generation through Diffusion-based Triplane Re-posing
di: Sheung, Eddie Pokming, et al.
Pubblicazione: (2025)
di: Sheung, Eddie Pokming, et al.
Pubblicazione: (2025)
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
di: Wang, Xingrui, et al.
Pubblicazione: (2024)
di: Wang, Xingrui, et al.
Pubblicazione: (2024)
DINeMo: Learning Neural Mesh Models with no 3D Annotations
di: Guo, Weijie, et al.
Pubblicazione: (2025)
di: Guo, Weijie, et al.
Pubblicazione: (2025)
4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
di: Zhong, Shanshan, et al.
Pubblicazione: (2025)
di: Zhong, Shanshan, et al.
Pubblicazione: (2025)
PASR: Pose-Aware 3D Shape Retrieval from Occluded Single Views
di: Shi, Jiaxin, et al.
Pubblicazione: (2026)
di: Shi, Jiaxin, et al.
Pubblicazione: (2026)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
di: Ma, Wufei, et al.
Pubblicazione: (2025)
di: Ma, Wufei, et al.
Pubblicazione: (2025)
DIRECT-3D: Learning Direct Text-to-3D Generation on Massive Noisy 3D Data
di: Liu, Qihao, et al.
Pubblicazione: (2024)
di: Liu, Qihao, et al.
Pubblicazione: (2024)
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
di: Liu, Qihao, et al.
Pubblicazione: (2025)
di: Liu, Qihao, et al.
Pubblicazione: (2025)
LychSim: A Controllable and Interactive Simulation Framework for Vision Research
di: Ma, Wufei, et al.
Pubblicazione: (2026)
di: Ma, Wufei, et al.
Pubblicazione: (2026)
NOVUM: Neural Object Volumes for Robust Object Classification
di: Jesslen, Artur, et al.
Pubblicazione: (2023)
di: Jesslen, Artur, et al.
Pubblicazione: (2023)
SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
di: Ma, Wufei, et al.
Pubblicazione: (2025)
di: Ma, Wufei, et al.
Pubblicazione: (2025)
SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference
di: Wang, Feng, et al.
Pubblicazione: (2023)
di: Wang, Feng, et al.
Pubblicazione: (2023)
Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution
di: Liu, Qihao, et al.
Pubblicazione: (2024)
di: Liu, Qihao, et al.
Pubblicazione: (2024)
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
di: Wang, Xingrui, et al.
Pubblicazione: (2025)
di: Wang, Xingrui, et al.
Pubblicazione: (2025)
3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
di: Ma, Wufei, et al.
Pubblicazione: (2024)
di: Ma, Wufei, et al.
Pubblicazione: (2024)
ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
di: Yang, Timing, et al.
Pubblicazione: (2025)
di: Yang, Timing, et al.
Pubblicazione: (2025)
Generating Images with 3D Annotations Using Diffusion Models
di: Ma, Wufei, et al.
Pubblicazione: (2023)
di: Ma, Wufei, et al.
Pubblicazione: (2023)
Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
di: Liu, Qihao, et al.
Pubblicazione: (2025)
di: Liu, Qihao, et al.
Pubblicazione: (2025)
TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
di: Zhang, Guofeng, et al.
Pubblicazione: (2025)
di: Zhang, Guofeng, et al.
Pubblicazione: (2025)
GenCA: A Text-conditioned Generative Model for Realistic and Drivable Codec Avatars
di: Sun, Keqiang, et al.
Pubblicazione: (2024)
di: Sun, Keqiang, et al.
Pubblicazione: (2024)
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
di: Ning, Zhenyu, et al.
Pubblicazione: (2025)
di: Ning, Zhenyu, et al.
Pubblicazione: (2025)
FullTransNet: Full Transformer with Local-Global Attention for Video Summarization
di: Lan, Libin, et al.
Pubblicazione: (2025)
di: Lan, Libin, et al.
Pubblicazione: (2025)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
di: Yuan, Huaying, et al.
Pubblicazione: (2025)
di: Yuan, Huaying, et al.
Pubblicazione: (2025)
Latent Dynamics for Full Body Avatar Animation
di: Peng, Shichong, et al.
Pubblicazione: (2026)
di: Peng, Shichong, et al.
Pubblicazione: (2026)
VideoAuteur: Towards Long Narrative Video Generation
di: Xiao, Junfei, et al.
Pubblicazione: (2025)
di: Xiao, Junfei, et al.
Pubblicazione: (2025)
Understanding Pan-Sharpening via Generalized Inverse
di: Liu, Shiqi, et al.
Pubblicazione: (2023)
di: Liu, Shiqi, et al.
Pubblicazione: (2023)
Computer Vision and Its Relationship to Cognitive Science: A perspective from Bayes Decision Theory
di: Yuille, Alan, et al.
Pubblicazione: (2026)
di: Yuille, Alan, et al.
Pubblicazione: (2026)
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
di: Xue, Zhucun, et al.
Pubblicazione: (2025)
di: Xue, Zhucun, et al.
Pubblicazione: (2025)
HECTOR: Hybrid Editable Compositional Object References for Video Generation
di: Zhang, Guofeng, et al.
Pubblicazione: (2026)
di: Zhang, Guofeng, et al.
Pubblicazione: (2026)
Thinking with Spatial Code for Physical-World Video Reasoning
di: Chen, Jieneng, et al.
Pubblicazione: (2026)
di: Chen, Jieneng, et al.
Pubblicazione: (2026)
E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation
di: Xu, Zeyu, et al.
Pubblicazione: (2025)
di: Xu, Zeyu, et al.
Pubblicazione: (2025)
Foundation VAEs for 3D CT Reconstruction, Augmentation, and Generation
di: Chen, Qi, et al.
Pubblicazione: (2026)
di: Chen, Qi, et al.
Pubblicazione: (2026)
Uncertainty-Aware Deep Video Compression with Ensembles
di: Ma, Wufei, et al.
Pubblicazione: (2024)
di: Ma, Wufei, et al.
Pubblicazione: (2024)
Rethinking Radiology Report Generation via Causal Inspired Counterfactual Augmentation
di: Song, Xiao, et al.
Pubblicazione: (2023)
di: Song, Xiao, et al.
Pubblicazione: (2023)
Video-Oasis: Rethinking Evaluation of Video Understanding
di: Lim, Geuntaek, et al.
Pubblicazione: (2026)
di: Lim, Geuntaek, et al.
Pubblicazione: (2026)
Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval
di: Liu, Weijia, et al.
Pubblicazione: (2025)
di: Liu, Weijia, et al.
Pubblicazione: (2025)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
di: Shen, Xiaoqian, et al.
Pubblicazione: (2025)
di: Shen, Xiaoqian, et al.
Pubblicazione: (2025)
VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs
di: Liao, Ruotong, et al.
Pubblicazione: (2024)
di: Liao, Ruotong, et al.
Pubblicazione: (2024)
DrVideo: Document Retrieval Based Long Video Understanding
di: Ma, Ziyu, et al.
Pubblicazione: (2024)
di: Ma, Ziyu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ImageNet3D: Towards General-Purpose Object-Level 3D Understanding
di: Ma, Wufei, et al.
Pubblicazione: (2024) -
TriDiff-4D: Fast 4D Generation through Diffusion-based Triplane Re-posing
di: Sheung, Eddie Pokming, et al.
Pubblicazione: (2025) -
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
di: Wang, Xingrui, et al.
Pubblicazione: (2024) -
DINeMo: Learning Neural Mesh Models with no 3D Annotations
di: Guo, Weijie, et al.
Pubblicazione: (2025) -
4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
di: Zhong, Shanshan, et al.
Pubblicazione: (2025)