LoCoNet: Long-Short Context Network for Active Speaker Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xizi, Cheng, Feng, Bertasius, Gedas, Crandall, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TimeRefine: Temporal Grounding with Time Refining Video LLM
von: Wang, Xizi, et al.
Veröffentlicht: (2024)
von: Wang, Xizi, et al.
Veröffentlicht: (2024)
RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos
von: Hannan, Tanveer, et al.
Veröffentlicht: (2023)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2023)
Sequence-Based Identification of First-Person Camera Wearers in Third-Person Views
von: Zhao, Ziwei, et al.
Veröffentlicht: (2025)
von: Zhao, Ziwei, et al.
Veröffentlicht: (2025)
BASKET: A Large-Scale Video Dataset for Fine-Grained Skill Estimation
von: Pan, Yulu, et al.
Veröffentlicht: (2025)
von: Pan, Yulu, et al.
Veröffentlicht: (2025)
BIMBA: Selective-Scan Compression for Long-Range Video Question Answering
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2025)
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2025)
EgoVIS@CVPR: PAIR-Net: Enhancing Egocentric Speaker Detection via Pretrained Audio-Visual Fusion and Alignment Loss
von: Wang, Yu, et al.
Veröffentlicht: (2025)
von: Wang, Yu, et al.
Veröffentlicht: (2025)
GateFusion: Hierarchical Gated Cross-Modal Fusion for Active Speaker Detection
von: Wang, Yu, et al.
Veröffentlicht: (2025)
von: Wang, Yu, et al.
Veröffentlicht: (2025)
Siamese Vision Transformers are Scalable Audio-visual Learners
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2024)
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2024)
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
von: Hannan, Tanveer, et al.
Veröffentlicht: (2024)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2024)
Video ReCap: Recursive Captioning of Hour-Long Videos
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2024)
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2024)
BOSS: Benchmark for Observation Space Shift in Long-Horizon Task
von: Yang, Yue, et al.
Veröffentlicht: (2025)
von: Yang, Yue, et al.
Veröffentlicht: (2025)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
A Simple LLM Framework for Long-Range Video Question-Answering
von: Zhang, Ce, et al.
Veröffentlicht: (2023)
von: Zhang, Ce, et al.
Veröffentlicht: (2023)
LiLo-VLA: Compositional Long-Horizon Manipulation via Linked Object-Centric Policies
von: Yang, Yue, et al.
Veröffentlicht: (2026)
von: Yang, Yue, et al.
Veröffentlicht: (2026)
TeDiO: Temporal Diagonal Optimization for Training-Free Coherent Video Diffusion
von: Tursynbek, Nurislam, et al.
Veröffentlicht: (2026)
von: Tursynbek, Nurislam, et al.
Veröffentlicht: (2026)
SiLVR: A Simple Language-based Video Reasoning Framework
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
DocSLM: A Small Vision-Language Model for Long Multimodal Document Understanding
von: Hannan, Tanveer, et al.
Veröffentlicht: (2025)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2025)
DAM: Dynamic Adapter Merging for Continual Video QA Learning
von: Cheng, Feng, et al.
Veröffentlicht: (2024)
von: Cheng, Feng, et al.
Veröffentlicht: (2024)
ExAct: A Video-Language Benchmark for Expert Action Analysis
von: Yi, Han, et al.
Veröffentlicht: (2025)
von: Yi, Han, et al.
Veröffentlicht: (2025)
Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2024)
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2024)
SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence
von: Pan, Yulu, et al.
Veröffentlicht: (2026)
von: Pan, Yulu, et al.
Veröffentlicht: (2026)
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2024)
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2024)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2026)
von: Wang, Ziyang, et al.
Veröffentlicht: (2026)
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
von: Zhou, Yiyang, et al.
Veröffentlicht: (2025)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2025)
TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs
von: Li, Baiqi, et al.
Veröffentlicht: (2026)
von: Li, Baiqi, et al.
Veröffentlicht: (2026)
ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video Synthesis
von: Fang, Yu, et al.
Veröffentlicht: (2025)
von: Fang, Yu, et al.
Veröffentlicht: (2025)
LoSh: Long-Short Text Joint Prediction Network for Referring Video Object Segmentation
von: Yuan, Linfeng, et al.
Veröffentlicht: (2023)
von: Yuan, Linfeng, et al.
Veröffentlicht: (2023)
HCF-Net: Hierarchical Context Fusion Network for Infrared Small Object Detection
von: Xu, Shibiao, et al.
Veröffentlicht: (2024)
von: Xu, Shibiao, et al.
Veröffentlicht: (2024)
ASDnB: Merging Face with Body Cues For Robust Active Speaker Detection
von: Roxo, Tiago, et al.
Veröffentlicht: (2024)
von: Roxo, Tiago, et al.
Veröffentlicht: (2024)
LoViC: Efficient Long Video Generation with Context Compression
von: Jiang, Jiaxiu, et al.
Veröffentlicht: (2025)
von: Jiang, Jiaxiu, et al.
Veröffentlicht: (2025)
MSCA-Net:Multi-Scale Context Aggregation Network for Infrared Small Target Detection
von: Lu, Xiaojin, et al.
Veröffentlicht: (2025)
von: Lu, Xiaojin, et al.
Veröffentlicht: (2025)
LoSA: Long-Short-range Adapter for Scaling End-to-End Temporal Action Localization
von: Gupta, Akshita, et al.
Veröffentlicht: (2024)
von: Gupta, Akshita, et al.
Veröffentlicht: (2024)
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2026)
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2026)
FriendNet: Detection-Friendly Dehazing Network
von: Fan, Yihua, et al.
Veröffentlicht: (2024)
von: Fan, Yihua, et al.
Veröffentlicht: (2024)
DF-Net: The Digital Forensics Network for Image Forgery Detection
von: Fischinger, David, et al.
Veröffentlicht: (2025)
von: Fischinger, David, et al.
Veröffentlicht: (2025)
In-Context LoRA for Diffusion Transformers
von: Huang, Lianghua, et al.
Veröffentlicht: (2024)
von: Huang, Lianghua, et al.
Veröffentlicht: (2024)
BIAS: A Body-based Interpretable Active Speaker Approach
von: Roxo, Tiago, et al.
Veröffentlicht: (2024)
von: Roxo, Tiago, et al.
Veröffentlicht: (2024)
OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels
von: Lou, Meng, et al.
Veröffentlicht: (2025)
von: Lou, Meng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TimeRefine: Temporal Grounding with Time Refining Video LLM
von: Wang, Xizi, et al.
Veröffentlicht: (2024) -
RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos
von: Hannan, Tanveer, et al.
Veröffentlicht: (2023) -
Sequence-Based Identification of First-Person Camera Wearers in Third-Person Views
von: Zhao, Ziwei, et al.
Veröffentlicht: (2025) -
BASKET: A Large-Scale Video Dataset for Fine-Grained Skill Estimation
von: Pan, Yulu, et al.
Veröffentlicht: (2025) -
BIMBA: Selective-Scan Compression for Long-Range Video Question Answering
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2025)