VQToken: Neural Discrete Token Representation Learning for Extreme Token Reduction in Video Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Haichao, Fu, Yun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dense Video Understanding with Gated Residual Tokenization
por: Zhang, Haichao, et al.
Publicado: (2025)
por: Zhang, Haichao, et al.
Publicado: (2025)
Person detection and re-identification in open-world settings of retail stores and public spaces
por: Brkljač, Branko, et al.
Publicado: (2025)
por: Brkljač, Branko, et al.
Publicado: (2025)
Deep Learning-Based Multi-Object Tracking: A Comprehensive Survey from Foundations to State-of-the-Art
por: Adžemović, Momir
Publicado: (2025)
por: Adžemović, Momir
Publicado: (2025)
MVTamperBench: Evaluating Robustness of Vision-Language Models
por: Agarwal, Amit, et al.
Publicado: (2024)
por: Agarwal, Amit, et al.
Publicado: (2024)
Transforming faces into video stories -- VideoFace2.0
por: Brkljač, Branko, et al.
Publicado: (2025)
por: Brkljač, Branko, et al.
Publicado: (2025)
Enhancing Maritime Object Detection in Real-Time with RT-DETR and Data Augmentation
por: Nemati, Nader
Publicado: (2025)
por: Nemati, Nader
Publicado: (2025)
Goal-Oriented Source Coding using LDPC Codes for Compressed-Domain Image Classification
por: Aliouat, Ahcen, et al.
Publicado: (2025)
por: Aliouat, Ahcen, et al.
Publicado: (2025)
A Practical Synthesis of Detecting AI-Generated Textual, Visual, and Audio Content
por: Cao, Lele
Publicado: (2025)
por: Cao, Lele
Publicado: (2025)
ISO/IEC-Compliant Match-on-Card Face Verification with Short Binary Templates
por: Ganmati, Abdelilah, et al.
Publicado: (2025)
por: Ganmati, Abdelilah, et al.
Publicado: (2025)
Cepstral Smoothing of Binary Masks for Convolutive Blind Separation of Speech Mixtures
por: Missaoui, Ibrahim, et al.
Publicado: (2026)
por: Missaoui, Ibrahim, et al.
Publicado: (2026)
Evaluating the Effect of Compression on Video Temporal Consistency Using Objective Quality Metrics
por: Zsoldos, Peter
Publicado: (2026)
por: Zsoldos, Peter
Publicado: (2026)
Enhancing rice leaf images: An overview of image denoising techniques
por: Chutia, Rupjyoti, et al.
Publicado: (2025)
por: Chutia, Rupjyoti, et al.
Publicado: (2025)
Probabilistic and nonlinear compressive sensing
por: Barth, Lukas Silvester, et al.
Publicado: (2025)
por: Barth, Lukas Silvester, et al.
Publicado: (2025)
Efficient and Robust Remote Sensing Image Denoising Using Randomized Approximation of Geodesics' Gramian on the Manifold Underlying the Patch Space
por: Gajamannage, Kelum, et al.
Publicado: (2025)
por: Gajamannage, Kelum, et al.
Publicado: (2025)
Texture Discrimination via Hilbert Curve Path Based Information Quantifiers
por: Bariviera, Aurelio F., et al.
Publicado: (2024)
por: Bariviera, Aurelio F., et al.
Publicado: (2024)
Attenuation-adjusted deep learning of pore defects in 2D radiographs of additive manufacturing powders
por: Bjerregaard, Andreas, et al.
Publicado: (2024)
por: Bjerregaard, Andreas, et al.
Publicado: (2024)
Eleven Primitives and Three Gates: The Universal Structure of Computational Imaging
por: Yang, Chengshuai, et al.
Publicado: (2026)
por: Yang, Chengshuai, et al.
Publicado: (2026)
IF-D: A High-Frequency, General-Purpose Inertial Foundation Dataset for Self-Supervised Learning
por: Ferreira, Patrick, et al.
Publicado: (2025)
por: Ferreira, Patrick, et al.
Publicado: (2025)
Image Denoising Using the Geodesics' Gramian of the Manifold Underlying Patch-Space
por: Gajamannage, Kelum
Publicado: (2020)
por: Gajamannage, Kelum
Publicado: (2020)
Efficient Image Denoising by Low-Rank Singular Vector Approximations of Geodesics' Gramian Matrix
por: Gajamannage, Kelum, et al.
Publicado: (2022)
por: Gajamannage, Kelum, et al.
Publicado: (2022)
Matlab-based Epoch Extraction for Speaker Differentiation
por: Li, Kunlun, et al.
Publicado: (2024)
por: Li, Kunlun, et al.
Publicado: (2024)
Detection of high-frequency oscillations using time-frequency analysis
por: Mohammadpour, Mostafa, et al.
Publicado: (2025)
por: Mohammadpour, Mostafa, et al.
Publicado: (2025)
TerraSeg: Self-Supervised Ground Segmentation for Any LiDAR
por: Lentsch, Ted, et al.
Publicado: (2026)
por: Lentsch, Ted, et al.
Publicado: (2026)
UNION: Unsupervised 3D Object Detection using Object Appearance-based Pseudo-Classes
por: Lentsch, Ted, et al.
Publicado: (2024)
por: Lentsch, Ted, et al.
Publicado: (2024)
STOPA: A Database of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and Attribution
por: Firc, Anton, et al.
Publicado: (2025)
por: Firc, Anton, et al.
Publicado: (2025)
ADD for Multi-Bit Image Watermarking
por: Luo, An, et al.
Publicado: (2026)
por: Luo, An, et al.
Publicado: (2026)
Understanding the Algorithm Behind Audio Key Detection
por: Silva, Henrique Perez G.
Publicado: (2025)
por: Silva, Henrique Perez G.
Publicado: (2025)
Segmenting the Complex and Irregular in Two-Phase Flows: A Real-World Empirical Study with SAM2
por: Küçük, Semanur, et al.
Publicado: (2025)
por: Küçük, Semanur, et al.
Publicado: (2025)
Proficiency-Aware Adaptation and Data Augmentation for Robust L2 ASR
por: Sun, Ling, et al.
Publicado: (2025)
por: Sun, Ling, et al.
Publicado: (2025)
Efficient Noise Calculation in Deep Learning-based MRI Reconstructions
por: Dalmaz, Onat, et al.
Publicado: (2025)
por: Dalmaz, Onat, et al.
Publicado: (2025)
Heart Failure Prediction using Modal Decomposition and Masked Autoencoders for Scarce Echocardiography Databases
por: Bell-Navas, Andrés, et al.
Publicado: (2025)
por: Bell-Navas, Andrés, et al.
Publicado: (2025)
Associative Syntax and Maximal Repetitions reveal context-dependent complexity in fruit bat communication
por: Assom, Luigi
Publicado: (2025)
por: Assom, Luigi
Publicado: (2025)
LinkedOut: Linking World Knowledge Representation Out of Video LLM for Next-Generation Video Recommendation
por: Zhang, Haichao, et al.
Publicado: (2025)
por: Zhang, Haichao, et al.
Publicado: (2025)
Out-of-Sight Embodied Agents: Multimodal Tracking, Sensor Fusion, and Trajectory Forecasting
por: Zhang, Haichao, et al.
Publicado: (2025)
por: Zhang, Haichao, et al.
Publicado: (2025)
The Voynich Codex Decoded: Statistical Symbolism and Scroll-Wide Logic
por: Jama, Suhaib A.
Publicado: (2025)
por: Jama, Suhaib A.
Publicado: (2025)
Designing Any Imaging System from Natural Language: Agent-Constrained Composition over a Finite Primitive Basis
por: Yang, Chengshuai
Publicado: (2026)
por: Yang, Chengshuai
Publicado: (2026)
HELIX: Scaling Raw Audio Understanding with Hybrid Mamba-Attention Beyond the Quadratic Limit
por: Khushiyant, et al.
Publicado: (2026)
por: Khushiyant, et al.
Publicado: (2026)
Ultrahigh-Q chiral resonances empowered by multi-head attention deep learning
por: Zhang, Cong, et al.
Publicado: (2025)
por: Zhang, Cong, et al.
Publicado: (2025)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
por: Patel, Urjitkumar, et al.
Publicado: (2025)
por: Patel, Urjitkumar, et al.
Publicado: (2025)
Tighter Bounds on the Information Bottleneck with Application to Deep Learning
por: Weingarten, Nir, et al.
Publicado: (2024)
por: Weingarten, Nir, et al.
Publicado: (2024)
Ejemplares similares
-
Dense Video Understanding with Gated Residual Tokenization
por: Zhang, Haichao, et al.
Publicado: (2025) -
Person detection and re-identification in open-world settings of retail stores and public spaces
por: Brkljač, Branko, et al.
Publicado: (2025) -
Deep Learning-Based Multi-Object Tracking: A Comprehensive Survey from Foundations to State-of-the-Art
por: Adžemović, Momir
Publicado: (2025) -
MVTamperBench: Evaluating Robustness of Vision-Language Models
por: Agarwal, Amit, et al.
Publicado: (2024) -
Transforming faces into video stories -- VideoFace2.0
por: Brkljač, Branko, et al.
Publicado: (2025)