Enregistré dans:
| Auteurs principaux: | Tang, Guowei, Qian, Tianwen, Zheng, Huanran, Wang, Yifei, Wang, Xiaoling |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2603.21493 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
par: Wang, Yifei, et autres
Publié: (2025)
par: Wang, Yifei, et autres
Publié: (2025)
SSNVC: Single Stream Neural Video Compression with Implicit Temporal Information
par: Wang, Feng, et autres
Publié: (2024)
par: Wang, Feng, et autres
Publié: (2024)
A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming
par: Zhou, Pengyuan, et autres
Publié: (2024)
par: Zhou, Pengyuan, et autres
Publié: (2024)
Looking Backward: Streaming Video-to-Video Translation with Feature Banks
par: Liang, Feng, et autres
Publié: (2024)
par: Liang, Feng, et autres
Publié: (2024)
Adaptive 3D Gaussian Splatting Video Streaming
par: Gong, Han, et autres
Publié: (2025)
par: Gong, Han, et autres
Publié: (2025)
Viewport Prediction for Volumetric Video Streaming by Exploring Video Saliency and Trajectory Information
par: Li, Jie, et autres
Publié: (2023)
par: Li, Jie, et autres
Publié: (2023)
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
par: Li, Chunyu, et autres
Publié: (2026)
par: Li, Chunyu, et autres
Publié: (2026)
A Multimodal Transformer for Live Streaming Highlight Prediction
par: Deng, Jiaxin, et autres
Publié: (2024)
par: Deng, Jiaxin, et autres
Publié: (2024)
SwinGS: Sliding Window Gaussian Splatting for Volumetric Video Streaming with Arbitrary Length
par: Liu, Bangya, et autres
Publié: (2024)
par: Liu, Bangya, et autres
Publié: (2024)
High-Quality Live Video Streaming via Transcoding Time Prediction and Preset Selection
par: Shahre-Babak, Zahra Nabizadeh, et autres
Publié: (2023)
par: Shahre-Babak, Zahra Nabizadeh, et autres
Publié: (2023)
Learning Long-Range Action Representation by Two-Stream Mamba Pyramid Network for Figure Skating Assessment
par: Wang, Fengshun, et autres
Publié: (2025)
par: Wang, Fengshun, et autres
Publié: (2025)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
par: Yuan, Huaying, et autres
Publié: (2025)
par: Yuan, Huaying, et autres
Publié: (2025)
Editing Physiological Signals in Videos Using Latent Representations
par: Zhou, Tianwen, et autres
Publié: (2025)
par: Zhou, Tianwen, et autres
Publié: (2025)
Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding
par: Huang, Dawei, et autres
Publié: (2025)
par: Huang, Dawei, et autres
Publié: (2025)
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
par: Kong, Fanheng, et autres
Publié: (2025)
par: Kong, Fanheng, et autres
Publié: (2025)
Scalable Event-Based Video Streaming for Machines with MoQ
par: Freeman, Andrew C.
Publié: (2025)
par: Freeman, Andrew C.
Publié: (2025)
LapisGS: Layered Progressive 3D Gaussian Splatting for Adaptive Streaming
par: Shi, Yuang, et autres
Publié: (2024)
par: Shi, Yuang, et autres
Publié: (2024)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
par: Zhang, Pingping, et autres
Publié: (2024)
par: Zhang, Pingping, et autres
Publié: (2024)
EVAN: Evolutional Video Streaming Adaptation via Neural Representation
par: Liu, Mufan, et autres
Publié: (2024)
par: Liu, Mufan, et autres
Publié: (2024)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
par: Xiao, Junbin, et autres
Publié: (2026)
par: Xiao, Junbin, et autres
Publié: (2026)
TMFNet: Two-Stream Multi-Channels Fusion Networks for Color Image Operation Chain Detection
par: Niu, Yakun, et autres
Publié: (2024)
par: Niu, Yakun, et autres
Publié: (2024)
Multimodal Fake News Video Explanation: Dataset, Analysis and Evaluation
par: Chen, Lizhi, et autres
Publié: (2025)
par: Chen, Lizhi, et autres
Publié: (2025)
VIoTGPT: Learning to Schedule Vision Tools in LLMs towards Intelligent Video Internet of Things
par: Zhong, Yaoyao, et autres
Publié: (2023)
par: Zhong, Yaoyao, et autres
Publié: (2023)
SequencePAR: Understanding Pedestrian Attributes via A Sequence Generation Paradigm
par: Jin, Jiandong, et autres
Publié: (2023)
par: Jin, Jiandong, et autres
Publié: (2023)
FMNV: A Dataset of Media-Published News Videos for Fake News Detection
par: Wang, Yihao, et autres
Publié: (2025)
par: Wang, Yihao, et autres
Publié: (2025)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
par: Wang, Jiapeng, et autres
Publié: (2024)
par: Wang, Jiapeng, et autres
Publié: (2024)
JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
par: Liu, Kai, et autres
Publié: (2026)
par: Liu, Kai, et autres
Publié: (2026)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
par: Wang, Shaoguang, et autres
Publié: (2026)
par: Wang, Shaoguang, et autres
Publié: (2026)
Error Analyses of Auto-Regressive Video Diffusion Models: A Unified Framework
par: Wang, Jing, et autres
Publié: (2025)
par: Wang, Jing, et autres
Publié: (2025)
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
par: Zhang, Chen-Lin, et autres
Publié: (2025)
par: Zhang, Chen-Lin, et autres
Publié: (2025)
Generative Frame Sampler for Long Video Understanding
par: Yao, Linli, et autres
Publié: (2025)
par: Yao, Linli, et autres
Publié: (2025)
Towards Realistic Low-Light Image Enhancement via ISP Driven Data Modeling
par: Wang, Zhihua, et autres
Publié: (2025)
par: Wang, Zhihua, et autres
Publié: (2025)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
par: Han, ZhaoYang, et autres
Publié: (2025)
par: Han, ZhaoYang, et autres
Publié: (2025)
Multi-Modal Image Fusion via Intervention-Stable Feature Learning
par: Wang, Xue, et autres
Publié: (2026)
par: Wang, Xue, et autres
Publié: (2026)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
par: Fang, Xinyu, et autres
Publié: (2024)
par: Fang, Xinyu, et autres
Publié: (2024)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
par: Wang, Yun, et autres
Publié: (2025)
par: Wang, Yun, et autres
Publié: (2025)
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
par: Wang, Xiang, et autres
Publié: (2025)
par: Wang, Xiang, et autres
Publié: (2025)
Towards Event-oriented Long Video Understanding
par: Du, Yifan, et autres
Publié: (2024)
par: Du, Yifan, et autres
Publié: (2024)
MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation
par: Shi, Haoyuan, et autres
Publié: (2026)
par: Shi, Haoyuan, et autres
Publié: (2026)
Color When It Counts: Grayscale-Guided Online Triggering for Always-On Streaming Video Sensing
par: Cai, Weitong, et autres
Publié: (2026)
par: Cai, Weitong, et autres
Publié: (2026)
Documents similaires
-
StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
par: Wang, Yifei, et autres
Publié: (2025) -
SSNVC: Single Stream Neural Video Compression with Implicit Temporal Information
par: Wang, Feng, et autres
Publié: (2024) -
A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming
par: Zhou, Pengyuan, et autres
Publié: (2024) -
Looking Backward: Streaming Video-to-Video Translation with Feature Banks
par: Liang, Feng, et autres
Publié: (2024) -
Adaptive 3D Gaussian Splatting Video Streaming
par: Gong, Han, et autres
Publié: (2025)