LumiVideo: An Intelligent Agentic System for Video Color Grading
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Yuchen, Gong, Junli, Cai, Hongmin, Cheung, Yiu-ming, Su, Weifeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Segmentation Models Understand the World? Towards Proactive Affordance Reasoning via Visual Chain-of-Thought
von: Guo, Yuchen, et al.
Veröffentlicht: (2026)
von: Guo, Yuchen, et al.
Veröffentlicht: (2026)
Fuse4Seg: Image Fusion for Multi-Modal Medical Segmentation via Bi-level Optimization
von: Guo, Yuchen, et al.
Veröffentlicht: (2024)
von: Guo, Yuchen, et al.
Veröffentlicht: (2024)
Adding Thermal Awareness to Visual Systems in Real-Time via Distilled Diffusion Models
von: Guo, Yuchen, et al.
Veröffentlicht: (2026)
von: Guo, Yuchen, et al.
Veröffentlicht: (2026)
PEAM: Parametric Embodied Agent Memory through Contrastive Internalization of Experience in Minecraft
von: Guo, Yuchen, et al.
Veröffentlicht: (2026)
von: Guo, Yuchen, et al.
Veröffentlicht: (2026)
Ask, Attend, Attack: A Effective Decision-Based Black-Box Targeted Attack for Image-to-Text Models
von: Zeng, Qingyuan, et al.
Veröffentlicht: (2024)
von: Zeng, Qingyuan, et al.
Veröffentlicht: (2024)
Agentic Video Intelligence: A Flexible Framework for Advanced Video Exploration and Understanding
von: Gao, Hong, et al.
Veröffentlicht: (2025)
von: Gao, Hong, et al.
Veröffentlicht: (2025)
Epistemic Uncertainty for Generated Image Detection
von: Nie, Jun, et al.
Veröffentlicht: (2024)
von: Nie, Jun, et al.
Veröffentlicht: (2024)
Adjusting Logit in Gaussian Form for Long-Tailed Visual Recognition
von: Li, Mengke, et al.
Veröffentlicht: (2023)
von: Li, Mengke, et al.
Veröffentlicht: (2023)
Preacher: Paper-to-Video Agentic System
von: Liu, Jingwei, et al.
Veröffentlicht: (2025)
von: Liu, Jingwei, et al.
Veröffentlicht: (2025)
DAE-Fuse: An Adaptive Discriminative Autoencoder for Multi-Modality Image Fusion
von: Guo, Yuchen, et al.
Veröffentlicht: (2024)
von: Guo, Yuchen, et al.
Veröffentlicht: (2024)
Bringing Multimodal Large Language Models to Infrared-Visible Image Fusion Quality Assessment
von: Guo, Yuchen, et al.
Veröffentlicht: (2026)
von: Guo, Yuchen, et al.
Veröffentlicht: (2026)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
von: Liu, Wenqi, et al.
Veröffentlicht: (2026)
von: Liu, Wenqi, et al.
Veröffentlicht: (2026)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
TYrPPG: Uncomplicated and Enhanced Learning Capability rPPG for Remote Heart Rate Estimation
von: Chen, Taixi, et al.
Veröffentlicht: (2025)
von: Chen, Taixi, et al.
Veröffentlicht: (2025)
LensWalk: Agentic Video Understanding by Planning How You See in Videos
von: Li, Keliang, et al.
Veröffentlicht: (2026)
von: Li, Keliang, et al.
Veröffentlicht: (2026)
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
LumiSculpt: Enabling Consistent Portrait Lighting in Video Generation
von: Zhang, Yuxin, et al.
Veröffentlicht: (2024)
von: Zhang, Yuxin, et al.
Veröffentlicht: (2024)
VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
von: Liang, Baoyu, et al.
Veröffentlicht: (2025)
von: Liang, Baoyu, et al.
Veröffentlicht: (2025)
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation
von: Zhao, Yiming, et al.
Veröffentlicht: (2026)
von: Zhao, Yiming, et al.
Veröffentlicht: (2026)
Prompt-based Consistent Video Colorization
von: Dani, Silvia, et al.
Veröffentlicht: (2025)
von: Dani, Silvia, et al.
Veröffentlicht: (2025)
VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority
von: Qiu, Chenhao, et al.
Veröffentlicht: (2026)
von: Qiu, Chenhao, et al.
Veröffentlicht: (2026)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning
von: Liu, Chengwen, et al.
Veröffentlicht: (2026)
von: Liu, Chengwen, et al.
Veröffentlicht: (2026)
AVA: Towards Agentic Video Analytics with Vision Language Models
von: Yan, Yuxuan, et al.
Veröffentlicht: (2025)
von: Yan, Yuxuan, et al.
Veröffentlicht: (2025)
GV-VAD : Exploring Video Generation for Weakly-Supervised Video Anomaly Detection
von: Cai, Suhang, et al.
Veröffentlicht: (2025)
von: Cai, Suhang, et al.
Veröffentlicht: (2025)
MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
von: Lin, Jingli, et al.
Veröffentlicht: (2025)
von: Lin, Jingli, et al.
Veröffentlicht: (2025)
Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation
von: Xiao, Zeqi, et al.
Veröffentlicht: (2025)
von: Xiao, Zeqi, et al.
Veröffentlicht: (2025)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
von: Yuan, Liping, et al.
Veröffentlicht: (2025)
von: Yuan, Liping, et al.
Veröffentlicht: (2025)
DiffusionAgent: Navigating Expert Models for Agentic Image Generation
von: Qin, Jie, et al.
Veröffentlicht: (2024)
von: Qin, Jie, et al.
Veröffentlicht: (2024)
Auto-US: An Ultrasound Video Diagnosis Agent Using Video Classification Framework and LLMs
von: Yang, Yuezhe, et al.
Veröffentlicht: (2025)
von: Yang, Yuezhe, et al.
Veröffentlicht: (2025)
A$^2$RD: Agentic Autoregressive Diffusion for Long Video Consistency
von: Long, Do Xuan, et al.
Veröffentlicht: (2026)
von: Long, Do Xuan, et al.
Veröffentlicht: (2026)
Video-T1: Test-Time Scaling for Video Generation
von: Liu, Fangfu, et al.
Veröffentlicht: (2025)
von: Liu, Fangfu, et al.
Veröffentlicht: (2025)
VideoMAR: Autoregressive Video Generatio with Continuous Tokens
von: Yu, Hu, et al.
Veröffentlicht: (2025)
von: Yu, Hu, et al.
Veröffentlicht: (2025)
Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
Think over Trajectories: Leveraging Video Generation to Reconstruct GPS Trajectories from Cellular Signaling
von: Zhang, Ruixing, et al.
Veröffentlicht: (2026)
von: Zhang, Ruixing, et al.
Veröffentlicht: (2026)
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks
von: Zhang, Han, et al.
Veröffentlicht: (2026)
von: Zhang, Han, et al.
Veröffentlicht: (2026)
Visual Agentic Memory: Enabling Online Long Video Understanding via Online Indexing, Hierarchical Memory, and Agentic Retrieval
von: Li, Aiden Yiliu, et al.
Veröffentlicht: (2026)
von: Li, Aiden Yiliu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Can Segmentation Models Understand the World? Towards Proactive Affordance Reasoning via Visual Chain-of-Thought
von: Guo, Yuchen, et al.
Veröffentlicht: (2026) -
Fuse4Seg: Image Fusion for Multi-Modal Medical Segmentation via Bi-level Optimization
von: Guo, Yuchen, et al.
Veröffentlicht: (2024) -
Adding Thermal Awareness to Visual Systems in Real-Time via Distilled Diffusion Models
von: Guo, Yuchen, et al.
Veröffentlicht: (2026) -
PEAM: Parametric Embodied Agent Memory through Contrastive Internalization of Experience in Minecraft
von: Guo, Yuchen, et al.
Veröffentlicht: (2026) -
Ask, Attend, Attack: A Effective Decision-Based Black-Box Targeted Attack for Image-to-Text Models
von: Zeng, Qingyuan, et al.
Veröffentlicht: (2024)