Guardado en:
| Autores principales: | Li, Haoxuan, Li, Mengyan, Zheng, Junjun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2601.07366 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HiVid: LLM-Guided Video Saliency For Content-Aware VOD And Live Streaming
por: Chen, Jiahui, et al.
Publicado: (2026)
por: Chen, Jiahui, et al.
Publicado: (2026)
RealCam-Vid: High-resolution Video Dataset with Dynamic Scenes and Metric-scale Camera Movements
por: Zheng, Guangcong, et al.
Publicado: (2025)
por: Zheng, Guangcong, et al.
Publicado: (2025)
HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation
por: Dong, Wenqi, et al.
Publicado: (2025)
por: Dong, Wenqi, et al.
Publicado: (2025)
HiDE: Hierarchical Dictionary-Based Entropy Modeling for Learned Image Compression
por: Xiong, Haoxuan, et al.
Publicado: (2026)
por: Xiong, Haoxuan, et al.
Publicado: (2026)
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
por: Lin, Kevin Qinghong, et al.
Publicado: (2025)
por: Lin, Kevin Qinghong, et al.
Publicado: (2025)
VideoAuteur: Towards Long Narrative Video Generation
por: Xiao, Junfei, et al.
Publicado: (2025)
por: Xiao, Junfei, et al.
Publicado: (2025)
HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising
por: Zou, Kai, et al.
Publicado: (2026)
por: Zou, Kai, et al.
Publicado: (2026)
LLMs Behind the Scenes: Enabling Narrative Scene Illustration
por: Roemmele, Melissa, et al.
Publicado: (2025)
por: Roemmele, Melissa, et al.
Publicado: (2025)
HiKER-SGG: Hierarchical Knowledge Enhanced Robust Scene Graph Generation
por: Zhang, Ce, et al.
Publicado: (2024)
por: Zhang, Ce, et al.
Publicado: (2024)
Orientation-anchored Hyper-Gaussian for 4D Reconstruction from Casual Videos
por: Wu, Junyi, et al.
Publicado: (2025)
por: Wu, Junyi, et al.
Publicado: (2025)
NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative
por: Nadeem, Asmar, et al.
Publicado: (2024)
por: Nadeem, Asmar, et al.
Publicado: (2024)
TextVidBench: A Benchmark for Long Video Scene Text Understanding
por: Zhong, Yangyang, et al.
Publicado: (2025)
por: Zhong, Yangyang, et al.
Publicado: (2025)
Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs
por: Ghorbani, Saeed
Publicado: (2025)
por: Ghorbani, Saeed
Publicado: (2025)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
por: Lan, Xiaohan, et al.
Publicado: (2024)
por: Lan, Xiaohan, et al.
Publicado: (2024)
BachVid: Training-Free Video Generation with Consistent Background and Character
por: Yan, Han, et al.
Publicado: (2025)
por: Yan, Han, et al.
Publicado: (2025)
HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
por: Meng, Yihao, et al.
Publicado: (2025)
por: Meng, Yihao, et al.
Publicado: (2025)
CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives
por: Meng, Yihao, et al.
Publicado: (2026)
por: Meng, Yihao, et al.
Publicado: (2026)
HiNeRV: Video Compression with Hierarchical Encoding-based Neural Representation
por: Kwan, Ho Man, et al.
Publicado: (2023)
por: Kwan, Ho Man, et al.
Publicado: (2023)
VidLeaks: Membership Inference Attacks Against Text-to-Video Models
por: Wang, Li, et al.
Publicado: (2026)
por: Wang, Li, et al.
Publicado: (2026)
HieraVid: Hierarchical Token Pruning for Fast Video Large Language Models
por: Guo, Yansong, et al.
Publicado: (2026)
por: Guo, Yansong, et al.
Publicado: (2026)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
por: Wang, Yi, et al.
Publicado: (2023)
por: Wang, Yi, et al.
Publicado: (2023)
Rethinking Autoregressive Models for Lossless Image Compression via Hierarchical Parallelism and Progressive Adaptation
por: Li, Daxin, et al.
Publicado: (2025)
por: Li, Daxin, et al.
Publicado: (2025)
HiGS: Hierarchical Generative Scene Framework for Multi-Step Associative Semantic Spatial Composition
por: Hong, Jiacheng, et al.
Publicado: (2025)
por: Hong, Jiacheng, et al.
Publicado: (2025)
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning
por: Jahagirdar, Soumya Shamarao, et al.
Publicado: (2026)
por: Jahagirdar, Soumya Shamarao, et al.
Publicado: (2026)
OmniVid: A Generative Framework for Universal Video Understanding
por: Wang, Junke, et al.
Publicado: (2024)
por: Wang, Junke, et al.
Publicado: (2024)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
por: Qin, Bosheng, et al.
Publicado: (2023)
por: Qin, Bosheng, et al.
Publicado: (2023)
STORYANCHORS: Generating Consistent Multi-Scene Story Frames for Long-Form Narratives
por: Wang, Bo, et al.
Publicado: (2025)
por: Wang, Bo, et al.
Publicado: (2025)
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
por: Wang, Xiaofeng, et al.
Publicado: (2024)
por: Wang, Xiaofeng, et al.
Publicado: (2024)
SmartDirector: Keyframe-Conditioned Cinematic Video Generation with Narrative Pacing Control
por: Zhang, Zhida, et al.
Publicado: (2026)
por: Zhang, Zhida, et al.
Publicado: (2026)
HiPrompt: Tuning-free Higher-Resolution Generation with Hierarchical MLLM Prompts
por: Liu, Xinyu, et al.
Publicado: (2024)
por: Liu, Xinyu, et al.
Publicado: (2024)
Narrative Aligned Long Form Video Question Answering
por: Jain, Rahul, et al.
Publicado: (2026)
por: Jain, Rahul, et al.
Publicado: (2026)
Vid3D: Synthesis of Dynamic 3D Scenes using 2D Video Diffusion
por: Parthasarathy, Rishab, et al.
Publicado: (2024)
por: Parthasarathy, Rishab, et al.
Publicado: (2024)
Generating Narrated Lecture Videos from Slides with Synchronized Highlights
por: Holmberg, Alexander
Publicado: (2025)
por: Holmberg, Alexander
Publicado: (2025)
STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot Narrative
por: Zhang, Peixuan, et al.
Publicado: (2025)
por: Zhang, Peixuan, et al.
Publicado: (2025)
MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval
por: Cai, Weitong, et al.
Publicado: (2024)
por: Cai, Weitong, et al.
Publicado: (2024)
MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives
por: Ikezogwo, Wisdom O., et al.
Publicado: (2025)
por: Ikezogwo, Wisdom O., et al.
Publicado: (2025)
Hi-GaTA: Hierarchical Gated Temporal Aggregation Adapter for Surgical Video Report Generation
por: Sun, Kedi, et al.
Publicado: (2026)
por: Sun, Kedi, et al.
Publicado: (2026)
VidCLearn: A Continual Learning Approach for Text-to-Video Generation
por: Zanchetta, Luca, et al.
Publicado: (2025)
por: Zanchetta, Luca, et al.
Publicado: (2025)
Controllable Generative Video Compression
por: Ding, Ding, et al.
Publicado: (2026)
por: Ding, Ding, et al.
Publicado: (2026)
SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
por: Li, Jungang, et al.
Publicado: (2024)
por: Li, Jungang, et al.
Publicado: (2024)
Ejemplares similares
-
HiVid: LLM-Guided Video Saliency For Content-Aware VOD And Live Streaming
por: Chen, Jiahui, et al.
Publicado: (2026) -
RealCam-Vid: High-resolution Video Dataset with Dynamic Scenes and Metric-scale Camera Movements
por: Zheng, Guangcong, et al.
Publicado: (2025) -
HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation
por: Dong, Wenqi, et al.
Publicado: (2025) -
HiDE: Hierarchical Dictionary-Based Entropy Modeling for Learned Image Compression
por: Xiong, Haoxuan, et al.
Publicado: (2026) -
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
por: Lin, Kevin Qinghong, et al.
Publicado: (2025)