Vript: A Video Is Worth Thousands of Words
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Dongjie, Huang, Suyuan, Lu, Chengqiang, Han, Xiaodong, Zhang, Haoxin, Gao, Yan, Hu, Yao, Zhao, Hai |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Video Is Not Worth a Thousand Words
por: Pollard, Sam, et al.
Publicado: (2025)
por: Pollard, Sam, et al.
Publicado: (2025)
From Image to Video, what do we need in multimodal LLMs?
por: Huang, Suyuan, et al.
Publicado: (2024)
por: Huang, Suyuan, et al.
Publicado: (2024)
WordVIS: A Color Worth A Thousand Words
por: Khan, Umar, et al.
Publicado: (2024)
por: Khan, Umar, et al.
Publicado: (2024)
Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image Generation
por: Tang, Raphael, et al.
Publicado: (2024)
por: Tang, Raphael, et al.
Publicado: (2024)
An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
por: Luo, Zhi, et al.
Publicado: (2025)
por: Luo, Zhi, et al.
Publicado: (2025)
Not Every Image is Worth a Thousand Words: Quantifying Originality in Stable Diffusion
por: Haviv, Adi, et al.
Publicado: (2024)
por: Haviv, Adi, et al.
Publicado: (2024)
One Image is Worth a Thousand Words: A Usability Preservable Text-Image Collaborative Erasing Framework
por: Li, Feiran, et al.
Publicado: (2025)
por: Li, Feiran, et al.
Publicado: (2025)
A LoRA is Worth a Thousand Pictures
por: Liu, Chenxi, et al.
Publicado: (2024)
por: Liu, Chenxi, et al.
Publicado: (2024)
Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models
por: Wang, Jiayu, et al.
Publicado: (2024)
por: Wang, Jiayu, et al.
Publicado: (2024)
Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation
por: Waseem, Faraz, et al.
Publicado: (2024)
por: Waseem, Faraz, et al.
Publicado: (2024)
Is a Picture Worth a Thousand Words? Adaptive Multimodal Fact-Checking with Visual Evidence Necessity
por: Jung, Jaeyoon, et al.
Publicado: (2026)
por: Jung, Jaeyoon, et al.
Publicado: (2026)
A Picture is Worth a Thousand Words? An Empirical Study of Aggregation Strategies for Visual Financial Document Retrieval
por: Lim, Ho Hung, et al.
Publicado: (2026)
por: Lim, Ho Hung, et al.
Publicado: (2026)
A Label is Worth a Thousand Images in Dataset Distillation
por: Qin, Tian, et al.
Publicado: (2024)
por: Qin, Tian, et al.
Publicado: (2024)
An Embedding is Worth a Thousand Noisy Labels
por: Di Salvo, Francesco, et al.
Publicado: (2024)
por: Di Salvo, Francesco, et al.
Publicado: (2024)
A Video is Worth 10,000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval
por: Gwilliam, Matthew, et al.
Publicado: (2023)
por: Gwilliam, Matthew, et al.
Publicado: (2023)
Fair-Eye Net: A Fair, Trustworthy, Multimodal Integrated Glaucoma Full Chain AI System
por: Wei, Wenbin, et al.
Publicado: (2026)
por: Wei, Wenbin, et al.
Publicado: (2026)
What is Point Supervision Worth in Video Instance Segmentation?
por: Huang, Shuaiyi, et al.
Publicado: (2024)
por: Huang, Shuaiyi, et al.
Publicado: (2024)
VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
por: Chen, Haoxin, et al.
Publicado: (2024)
por: Chen, Haoxin, et al.
Publicado: (2024)
A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image Features
por: Barroso-Laguna, Axel, et al.
Publicado: (2025)
por: Barroso-Laguna, Axel, et al.
Publicado: (2025)
Malware Detection in Docker Containers: An Image is Worth a Thousand Logs
por: Nousias, Akis, et al.
Publicado: (2025)
por: Nousias, Akis, et al.
Publicado: (2025)
EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
por: Liu, Yaofang, et al.
Publicado: (2023)
por: Liu, Yaofang, et al.
Publicado: (2023)
A Task is Worth One Word: Learning with Task Prompts for High-Quality Versatile Image Inpainting
por: Zhuang, Junhao, et al.
Publicado: (2023)
por: Zhuang, Junhao, et al.
Publicado: (2023)
Learning Interaction-aware 3D Gaussian Splatting for One-shot Hand Avatars
por: Huang, Xuan, et al.
Publicado: (2024)
por: Huang, Xuan, et al.
Publicado: (2024)
Images are Worth Variable Length of Representations
por: Mao, Lingjun, et al.
Publicado: (2025)
por: Mao, Lingjun, et al.
Publicado: (2025)
A Video is Worth 256 Bases: Spatial-Temporal Expectation-Maximization Inversion for Zero-Shot Video Editing
por: Li, Maomao, et al.
Publicado: (2023)
por: Li, Maomao, et al.
Publicado: (2023)
ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition
por: Huang, Ronggang, et al.
Publicado: (2025)
por: Huang, Ronggang, et al.
Publicado: (2025)
A Thousand Words or An Image: Studying the Influence of Persona Modality in Multimodal LLMs
por: Broomfield, Julius, et al.
Publicado: (2025)
por: Broomfield, Julius, et al.
Publicado: (2025)
Mogo: RQ Hierarchical Causal Transformer for High-Quality 3D Human Motion Generation
por: Fu, Dongjie
Publicado: (2024)
por: Fu, Dongjie
Publicado: (2024)
Are a Thousand Words Better Than a Single Picture? Beyond Images -- A Framework for Multi-Modal Knowledge Graph Dataset Enrichment
por: Zhang, Pengyu, et al.
Publicado: (2026)
por: Zhang, Pengyu, et al.
Publicado: (2026)
Seeing a Rose in Five Thousand Ways
por: Zhang, Yunzhi, et al.
Publicado: (2022)
por: Zhang, Yunzhi, et al.
Publicado: (2022)
A Picture is Worth a Thousand (Correct) Captions: A Vision-Guided Judge-Corrector System for Multimodal Machine Translation
por: Betala, Siddharth, et al.
Publicado: (2025)
por: Betala, Siddharth, et al.
Publicado: (2025)
Repeating Words for Video-Language Retrieval with Coarse-to-Fine Objectives
por: Zhao, Haoyu, et al.
Publicado: (2025)
por: Zhao, Haoyu, et al.
Publicado: (2025)
A Picture is Worth a Thousand Prompts? Efficacy of Iterative Human-Driven Prompt Refinement in Image Regeneration Tasks
por: Trinh, Khoi, et al.
Publicado: (2025)
por: Trinh, Khoi, et al.
Publicado: (2025)
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models
por: Qu, Mengxue, et al.
Publicado: (2024)
por: Qu, Mengxue, et al.
Publicado: (2024)
Noise Calibration: Plug-and-play Content-Preserving Video Enhancement using Pre-trained Video Diffusion Models
por: Yang, Qinyu, et al.
Publicado: (2024)
por: Yang, Qinyu, et al.
Publicado: (2024)
iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models
por: Hu, Lianyu, et al.
Publicado: (2024)
por: Hu, Lianyu, et al.
Publicado: (2024)
CutClaw: Agentic Hours-Long Video Editing via Music Synchronization
por: Zhao, Shifang, et al.
Publicado: (2026)
por: Zhao, Shifang, et al.
Publicado: (2026)
VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
por: Zhao, Fufangchen, et al.
Publicado: (2025)
por: Zhao, Fufangchen, et al.
Publicado: (2025)
Learning to Animate Images from A Few Videos to Portray Delicate Human Actions
por: Li, Haoxin, et al.
Publicado: (2025)
por: Li, Haoxin, et al.
Publicado: (2025)
Towards Long-Form Spatio-Temporal Video Grounding
por: Gu, Xin, et al.
Publicado: (2026)
por: Gu, Xin, et al.
Publicado: (2026)
Ejemplares similares
-
A Video Is Not Worth a Thousand Words
por: Pollard, Sam, et al.
Publicado: (2025) -
From Image to Video, what do we need in multimodal LLMs?
por: Huang, Suyuan, et al.
Publicado: (2024) -
WordVIS: A Color Worth A Thousand Words
por: Khan, Umar, et al.
Publicado: (2024) -
Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image Generation
por: Tang, Raphael, et al.
Publicado: (2024) -
An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
por: Luo, Zhi, et al.
Publicado: (2025)