Dense Video Understanding with Gated Residual Tokenization
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Haichao, Chai, Wenhao, He, Shwai, Li, Ang, Fu, Yun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LinkedOut: Linking World Knowledge Representation Out of Video LLM for Next-Generation Video Recommendation
by: Zhang, Haichao, et al.
Published: (2025)
by: Zhang, Haichao, et al.
Published: (2025)
VQToken: Neural Discrete Token Representation Learning for Extreme Token Reduction in Video Large Language Models
by: Zhang, Haichao, et al.
Published: (2025)
by: Zhang, Haichao, et al.
Published: (2025)
TerraSeg: Self-Supervised Ground Segmentation for Any LiDAR
by: Lentsch, Ted, et al.
Published: (2026)
by: Lentsch, Ted, et al.
Published: (2026)
UNION: Unsupervised 3D Object Detection using Object Appearance-based Pseudo-Classes
by: Lentsch, Ted, et al.
Published: (2024)
by: Lentsch, Ted, et al.
Published: (2024)
A Practical Synthesis of Detecting AI-Generated Textual, Visual, and Audio Content
by: Cao, Lele
Published: (2025)
by: Cao, Lele
Published: (2025)
Deep Learning-Based Multi-Object Tracking: A Comprehensive Survey from Foundations to State-of-the-Art
by: Adžemović, Momir
Published: (2025)
by: Adžemović, Momir
Published: (2025)
Heart Failure Prediction using Modal Decomposition and Masked Autoencoders for Scarce Echocardiography Databases
by: Bell-Navas, Andrés, et al.
Published: (2025)
by: Bell-Navas, Andrés, et al.
Published: (2025)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
by: Patel, Urjitkumar, et al.
Published: (2025)
by: Patel, Urjitkumar, et al.
Published: (2025)
Benchmarking changepoint detection algorithms on cardiac time series
by: Cakmak, Ayse, et al.
Published: (2024)
by: Cakmak, Ayse, et al.
Published: (2024)
Named entity recognition for Serbian legal documents: Design, methodology and dataset development
by: Kalušev, Vladimir, et al.
Published: (2025)
by: Kalušev, Vladimir, et al.
Published: (2025)
Polarization-Based Eye Tracking with Personalized Siamese Architectures
by: Kalkanli, Beyza, et al.
Published: (2026)
by: Kalkanli, Beyza, et al.
Published: (2026)
Emerging-properties Mapping Using Spatial Embedding Statistics: EMUSES
by: Foulon, Chris, et al.
Published: (2024)
by: Foulon, Chris, et al.
Published: (2024)
Classifier Calibration at Scale: An Empirical Study of Model-Agnostic Post-Hoc Methods
by: Manokhin, Valery, et al.
Published: (2026)
by: Manokhin, Valery, et al.
Published: (2026)
AI-Enhanced Acoustic Analysis for Comprehensive Biodiversity Monitoring and Assessment
by: Bobba, Kumar Srinivas, et al.
Published: (2024)
by: Bobba, Kumar Srinivas, et al.
Published: (2024)
MVTamperBench: Evaluating Robustness of Vision-Language Models
by: Agarwal, Amit, et al.
Published: (2024)
by: Agarwal, Amit, et al.
Published: (2024)
Visualizing the Evolution of Twitter (X.com) Conversations: A Comprehensive Methodology Applied to AI Training Discussions on ChatGPT
by: Jess, Nicole, et al.
Published: (2024)
by: Jess, Nicole, et al.
Published: (2024)
Deep Spectral Meshes: Multi-Frequency Facial Mesh Processing with Graph Neural Networks
by: Kosk, Robert, et al.
Published: (2024)
by: Kosk, Robert, et al.
Published: (2024)
Optimizing Indoor Environmental Quality in Smart Buildings Using Deep Learning
by: Sabiri, Youssef, et al.
Published: (2025)
by: Sabiri, Youssef, et al.
Published: (2025)
MRMS-Net and LMRMS-Net: Scalable Multi-Representation Multi-Scale Networks for Time Series Classification
by: Alagöz, Celal, et al.
Published: (2026)
by: Alagöz, Celal, et al.
Published: (2026)
GeoJEPA: Towards Eliminating Augmentation- and Sampling Bias in Multimodal Geospatial Learning
by: Lundqvist, Theodor, et al.
Published: (2025)
by: Lundqvist, Theodor, et al.
Published: (2025)
Contract-Driven QoE Auditing for Speech and Singing Services: From MOS Regression to Service Graphs
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
The Influence of Color Stimuli on Adolescents' Emotion Playing Mobile Games
by: Kallabis, Leonie, et al.
Published: (2024)
by: Kallabis, Leonie, et al.
Published: (2024)
See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
by: Dong, Zixuan, et al.
Published: (2025)
by: Dong, Zixuan, et al.
Published: (2025)
Extraction Of Cumulative Blobs From Dynamic Gestures
by: Naulakha, Rishabh, et al.
Published: (2025)
by: Naulakha, Rishabh, et al.
Published: (2025)
Evaluating the Effect of Compression on Video Temporal Consistency Using Objective Quality Metrics
by: Zsoldos, Peter
Published: (2026)
by: Zsoldos, Peter
Published: (2026)
Can Agentic AI Match the Performance of Human Data Scientists?
by: Luo, An, et al.
Published: (2025)
by: Luo, An, et al.
Published: (2025)
AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science
by: Luo, An, et al.
Published: (2026)
by: Luo, An, et al.
Published: (2026)
Multimodal AI-based visualization of strategic leaders' emotional dynamics: a deep behavioral analysis of Trump's trade war discourse
by: Meng, Wei
Published: (2025)
by: Meng, Wei
Published: (2025)
Enhancing Maritime Object Detection in Real-Time with RT-DETR and Data Augmentation
by: Nemati, Nader
Published: (2025)
by: Nemati, Nader
Published: (2025)
A Tractography Analysis Framework Using Diffusion Maps to Study Thalamic Connectivity in Traumatic Brain Injury
by: Sharma, Akul, et al.
Published: (2025)
by: Sharma, Akul, et al.
Published: (2025)
SQuARE: Structured Query & Adaptive Retrieval Engine For Tabular Formats
by: Gondhalekar, Chinmay, et al.
Published: (2025)
by: Gondhalekar, Chinmay, et al.
Published: (2025)
DeepShade: Enable Shade Simulation by Text-conditioned Image Generation
by: Da, Longchao, et al.
Published: (2025)
by: Da, Longchao, et al.
Published: (2025)
Bridging Industrial Expertise and XR with LLM-Powered Conversational Agents
by: Tomkou, Despina, et al.
Published: (2025)
by: Tomkou, Despina, et al.
Published: (2025)
Accelerating Cerebral Diagnostics with BrainFusion: A Comprehensive MRI Tumor Framework
by: Houmaidi, Walid, et al.
Published: (2025)
by: Houmaidi, Walid, et al.
Published: (2025)
AttriGen: Automated Multi-Attribute Annotation for Blood Cell Datasets
by: Houmaidi, Walid, et al.
Published: (2025)
by: Houmaidi, Walid, et al.
Published: (2025)
EYE-DEX: Eye Disease Detection and EXplanation System
by: Sabiri, Youssef, et al.
Published: (2025)
by: Sabiri, Youssef, et al.
Published: (2025)
Causal Deep Learning
by: Vasilescu, M. Alex O.
Published: (2023)
by: Vasilescu, M. Alex O.
Published: (2023)
Graph Connectionist Temporal Classification for Phoneme Recognition
by: Grafé, Henry, et al.
Published: (2025)
by: Grafé, Henry, et al.
Published: (2025)
AssistedDS: Benchmarking How External Domain Knowledge Assists LLMs in Automated Data Science
by: Luo, An, et al.
Published: (2025)
by: Luo, An, et al.
Published: (2025)
Mitigating Catastrophic Forgetting in Streaming Generative and Predictive Learning via Stateful Replay
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
Similar Items
-
LinkedOut: Linking World Knowledge Representation Out of Video LLM for Next-Generation Video Recommendation
by: Zhang, Haichao, et al.
Published: (2025) -
VQToken: Neural Discrete Token Representation Learning for Extreme Token Reduction in Video Large Language Models
by: Zhang, Haichao, et al.
Published: (2025) -
TerraSeg: Self-Supervised Ground Segmentation for Any LiDAR
by: Lentsch, Ted, et al.
Published: (2026) -
UNION: Unsupervised 3D Object Detection using Object Appearance-based Pseudo-Classes
by: Lentsch, Ted, et al.
Published: (2024) -
A Practical Synthesis of Detecting AI-Generated Textual, Visual, and Audio Content
by: Cao, Lele
Published: (2025)