DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Yudong, Guo, Qingpei, Pan, Liyuan, Liu, Liu, Guan, Yu, Yang, Ming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
von: Han, Yudong, et al.
Veröffentlicht: (2026)
von: Han, Yudong, et al.
Veröffentlicht: (2026)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
von: Li, Huibin, et al.
Veröffentlicht: (2025)
von: Li, Huibin, et al.
Veröffentlicht: (2025)
A Simple Baseline for Streaming Video Understanding
von: Shen, Yujiao, et al.
Veröffentlicht: (2026)
von: Shen, Yujiao, et al.
Veröffentlicht: (2026)
SPMamba-YOLO: An Underwater Object Detection Network Based on Multi-Scale Feature Enhancement and Global Context Modeling
von: Liao, Guanghao, et al.
Veröffentlicht: (2026)
von: Liao, Guanghao, et al.
Veröffentlicht: (2026)
MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition
von: Hu, Yudong, et al.
Veröffentlicht: (2025)
von: Hu, Yudong, et al.
Veröffentlicht: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
NAC-TCN: Temporal Convolutional Networks with Causal Dilated Neighborhood Attention for Emotion Understanding
von: Mehta, Alexander, et al.
Veröffentlicht: (2023)
von: Mehta, Alexander, et al.
Veröffentlicht: (2023)
SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy
von: Elsharkawi, Ismael, et al.
Veröffentlicht: (2026)
von: Elsharkawi, Ismael, et al.
Veröffentlicht: (2026)
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
FoR-Net: Learning to Focus on Hard Regions for Efficient Semantic Segmentation
von: Chan, Sheng-Wei, et al.
Veröffentlicht: (2026)
von: Chan, Sheng-Wei, et al.
Veröffentlicht: (2026)
3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding
von: Chen, Yiping, et al.
Veröffentlicht: (2026)
von: Chen, Yiping, et al.
Veröffentlicht: (2026)
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
von: Feng, Yigui, et al.
Veröffentlicht: (2026)
von: Feng, Yigui, et al.
Veröffentlicht: (2026)
Cross-Modal Transfer from Memes to Videos: Addressing Data Scarcity in Hateful Video Detection
von: Wang, Han, et al.
Veröffentlicht: (2025)
von: Wang, Han, et al.
Veröffentlicht: (2025)
FocusedAD: Character-centric Movie Audio Description
von: Ye, Xiaojun, et al.
Veröffentlicht: (2025)
von: Ye, Xiaojun, et al.
Veröffentlicht: (2025)
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
von: He, Jianxiang, et al.
Veröffentlicht: (2025)
von: He, Jianxiang, et al.
Veröffentlicht: (2025)
A Reverse Causal Framework to Mitigate Spurious Correlations for Debiasing Scene Graph Generation
von: Sun, Shuzhou, et al.
Veröffentlicht: (2025)
von: Sun, Shuzhou, et al.
Veröffentlicht: (2025)
HY-Himmel Technical Report: Hierarchical Interleaved Multi-stream Motion Encoding for Long Video Understanding
von: Jin, Haopeng, et al.
Veröffentlicht: (2026)
von: Jin, Haopeng, et al.
Veröffentlicht: (2026)
Context-Aware Network Based on Multi-scale Spatio-temporal Attention for Action Recognition in Videos
von: Li, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyang, et al.
Veröffentlicht: (2025)
Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models
von: Dubois, L'ea, et al.
Veröffentlicht: (2025)
von: Dubois, L'ea, et al.
Veröffentlicht: (2025)
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
von: Li, Yayuan, et al.
Veröffentlicht: (2025)
von: Li, Yayuan, et al.
Veröffentlicht: (2025)
Foreground Focus: Enhancing Coherence and Fidelity in Camouflaged Image Generation
von: Chen, Pei-Chi, et al.
Veröffentlicht: (2025)
von: Chen, Pei-Chi, et al.
Veröffentlicht: (2025)
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
M3CAD: Towards Generic Cooperative Autonomous Driving Benchmark
von: Zhu, Morui, et al.
Veröffentlicht: (2025)
von: Zhu, Morui, et al.
Veröffentlicht: (2025)
Computer Vision for Clinical Gait Analysis: A Gait Abnormality Video Dataset
von: Ranjan, Rahm, et al.
Veröffentlicht: (2024)
von: Ranjan, Rahm, et al.
Veröffentlicht: (2024)
Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think
von: Tian, Jie, et al.
Veröffentlicht: (2025)
von: Tian, Jie, et al.
Veröffentlicht: (2025)
PhysVideoGenerator: Towards Physically Aware Video Generation via Latent Physics Guidance
von: Satish, Siddarth Nilol Kundur, et al.
Veröffentlicht: (2026)
von: Satish, Siddarth Nilol Kundur, et al.
Veröffentlicht: (2026)
Positive Style Accumulation: A Style Screening and Continuous Utilization Framework for Federated DG-ReID
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
Classifying Simulated Gait Impairments using Privacy-preserving Explainable Artificial Intelligence and Mobile Phone Videos
von: Reddy, Lauhitya, et al.
Veröffentlicht: (2024)
von: Reddy, Lauhitya, et al.
Veröffentlicht: (2024)
Textual and Visual Guided Task Adaptation for Source-Free Cross-Domain Few-Shot Segmentation
von: Liu, Jianming, et al.
Veröffentlicht: (2025)
von: Liu, Jianming, et al.
Veröffentlicht: (2025)
A Challenging Benchmark of Anime Style Recognition
von: Li, Haotang, et al.
Veröffentlicht: (2022)
von: Li, Haotang, et al.
Veröffentlicht: (2022)
Single-Shot Metric Depth from Focused Plenoptic Cameras
von: Lasheras-Hernandez, Blanca, et al.
Veröffentlicht: (2024)
von: Lasheras-Hernandez, Blanca, et al.
Veröffentlicht: (2024)
SemanticHuman-HD: High-Resolution Semantic Disentangled 3D Human Generation
von: Zheng, Peng, et al.
Veröffentlicht: (2024)
von: Zheng, Peng, et al.
Veröffentlicht: (2024)
Learning Association via Track-Detection Matching for Multi-Object Tracking
von: Adžemović, Momir
Veröffentlicht: (2025)
von: Adžemović, Momir
Veröffentlicht: (2025)
ERNet: Efficient Non-Rigid Registration Network for Point Sequences
von: He, Guangzhao, et al.
Veröffentlicht: (2025)
von: He, Guangzhao, et al.
Veröffentlicht: (2025)
Action Anticipation from SoccerNet Football Video Broadcasts
von: Dalal, Mohamad, et al.
Veröffentlicht: (2025)
von: Dalal, Mohamad, et al.
Veröffentlicht: (2025)
Hierarchical Feature-level Reverse Propagation for Post-Training Neural Networks
von: Ding, Ni, et al.
Veröffentlicht: (2025)
von: Ding, Ni, et al.
Veröffentlicht: (2025)
Sequence Matters: Harnessing Video Models in 3D Super-Resolution
von: Ko, Hyun-kyu, et al.
Veröffentlicht: (2024)
von: Ko, Hyun-kyu, et al.
Veröffentlicht: (2024)
Web-Scale Collection of Video Data for 4D Animal Reconstruction
von: Zhao, Brian Nlong, et al.
Veröffentlicht: (2025)
von: Zhao, Brian Nlong, et al.
Veröffentlicht: (2025)
RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
von: Qian, Wenxu, et al.
Veröffentlicht: (2025)
von: Qian, Wenxu, et al.
Veröffentlicht: (2025)
ViG-LRGC: Vision Graph Neural Networks with Learnable Reparameterized Graph Construction
von: Elsharkawi, Ismael, et al.
Veröffentlicht: (2025)
von: Elsharkawi, Ismael, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
von: Han, Yudong, et al.
Veröffentlicht: (2026) -
U-Net-Like Spiking Neural Networks for Single Image Dehazing
von: Li, Huibin, et al.
Veröffentlicht: (2025) -
A Simple Baseline for Streaming Video Understanding
von: Shen, Yujiao, et al.
Veröffentlicht: (2026) -
SPMamba-YOLO: An Underwater Object Detection Network Based on Multi-Scale Feature Enhancement and Global Context Modeling
von: Liao, Guanghao, et al.
Veröffentlicht: (2026) -
MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition
von: Hu, Yudong, et al.
Veröffentlicht: (2025)