V-LynX: Token Interface Alignment for Video+X LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Jungin, Lee, Jiyoung, Sohn, Kwanghoon |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization
by: Park, Jungin, et al.
Published: (2025)
by: Park, Jungin, et al.
Published: (2025)
Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations
by: Park, Jungin, et al.
Published: (2025)
by: Park, Jungin, et al.
Published: (2025)
Saliency-Aware Model Merging
by: Park, Jungin, et al.
Published: (2026)
by: Park, Jungin, et al.
Published: (2026)
Bridging Vision and Language Spaces with Assignment Prediction
by: Park, Jungin, et al.
Published: (2024)
by: Park, Jungin, et al.
Published: (2024)
Faster Parameter-Efficient Tuning with Token Redundancy Reduction
by: Kim, Kwonyoung, et al.
Published: (2025)
by: Kim, Kwonyoung, et al.
Published: (2025)
PointFix: Learning to Fix Domain Bias for Robust Online Stereo Adaptation
by: Kim, Kwonyoung, et al.
Published: (2022)
by: Kim, Kwonyoung, et al.
Published: (2022)
EBDM: Exemplar-guided Image Translation with Brownian-bridge Diffusion Models
by: Lee, Eungbean, et al.
Published: (2024)
by: Lee, Eungbean, et al.
Published: (2024)
Descriptive Image-Text Matching with Graded Contextual Similarity
by: Jang, Jinhyun, et al.
Published: (2025)
by: Jang, Jinhyun, et al.
Published: (2025)
Enhancing Source-Free Domain Adaptive Object Detection with Low-confidence Pseudo Label Distillation
by: Yoon, Ilhoon, et al.
Published: (2024)
by: Yoon, Ilhoon, et al.
Published: (2024)
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
by: Hyun, Jeongseok, et al.
Published: (2025)
by: Hyun, Jeongseok, et al.
Published: (2025)
Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions
by: Kwon, Jihoon, et al.
Published: (2025)
by: Kwon, Jihoon, et al.
Published: (2025)
Leveraging Textual Compositional Reasoning for Robust Change Captioning
by: Park, Kyu Ri, et al.
Published: (2025)
by: Park, Kyu Ri, et al.
Published: (2025)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
by: Qin, Jialong, et al.
Published: (2025)
by: Qin, Jialong, et al.
Published: (2025)
OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models
by: Kang, Minseok, et al.
Published: (2026)
by: Kang, Minseok, et al.
Published: (2026)
Rethinking Open-World Semi-Supervised Learning: Distribution Mismatch and Inductive Inference
by: Park, Seongheon, et al.
Published: (2024)
by: Park, Seongheon, et al.
Published: (2024)
X2SAM: Any Segmentation in Images and Videos
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
EditSplat: Multi-View Fusion and Attention-Guided Optimization for View-Consistent 3D Scene Editing with 3D Gaussian Splatting
by: Lee, Dong In, et al.
Published: (2024)
by: Lee, Dong In, et al.
Published: (2024)
VideoMAR: Autoregressive Video Generatio with Continuous Tokens
by: Yu, Hu, et al.
Published: (2025)
by: Yu, Hu, et al.
Published: (2025)
InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression
by: Lu, Dongchen, et al.
Published: (2025)
by: Lu, Dongchen, et al.
Published: (2025)
Plan-X: Instruct Video Generation via Semantic Planning
by: Huang, Lun, et al.
Published: (2025)
by: Huang, Lun, et al.
Published: (2025)
Knowledge-Guided Textual Reasoning for Explainable Video Anomaly Detection via LLMs
by: Lee, Hari
Published: (2025)
by: Lee, Hari
Published: (2025)
YTCommentQA: Video Question Answerability in Instructional Videos
by: Yang, Saelyne, et al.
Published: (2024)
by: Yang, Saelyne, et al.
Published: (2024)
TokenTrim: Inference-Time Token Pruning for Autoregressive Long Video Generation
by: Shaulov, Ariel, et al.
Published: (2026)
by: Shaulov, Ariel, et al.
Published: (2026)
Token-Level Inference-Time Alignment for Vision-Language Models
by: Chen, Kejia, et al.
Published: (2025)
by: Chen, Kejia, et al.
Published: (2025)
V.I.P. : Iterative Online Preference Distillation for Efficient Video Diffusion Models
by: Kim, Jisoo, et al.
Published: (2025)
by: Kim, Jisoo, et al.
Published: (2025)
LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation
by: Xing, Jiazheng, et al.
Published: (2026)
by: Xing, Jiazheng, et al.
Published: (2026)
GameGen-X: Interactive Open-world Game Video Generation
by: Che, Haoxuan, et al.
Published: (2024)
by: Che, Haoxuan, et al.
Published: (2024)
Spectral Motion Alignment for Video Motion Transfer using Diffusion Models
by: Park, Geon Yeong, et al.
Published: (2024)
by: Park, Geon Yeong, et al.
Published: (2024)
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
by: Lee, SuBeen, et al.
Published: (2025)
by: Lee, SuBeen, et al.
Published: (2025)
Generating Accurate and Detailed Captions for High-Resolution Images
by: Lee, Hankyeol, et al.
Published: (2025)
by: Lee, Hankyeol, et al.
Published: (2025)
AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding
by: Qi, Haozhe, et al.
Published: (2026)
by: Qi, Haozhe, et al.
Published: (2026)
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
by: Ji, Yicheng, et al.
Published: (2025)
by: Ji, Yicheng, et al.
Published: (2025)
Temporal Inversion for Learning Interval Change in Chest X-Rays
by: Ko, Hanbin, et al.
Published: (2026)
by: Ko, Hanbin, et al.
Published: (2026)
SToRM: Supervised Token Reduction for Multi-modal LLMs toward efficient end-to-end autonomous driving
by: Kim, Seo Hyun, et al.
Published: (2026)
by: Kim, Seo Hyun, et al.
Published: (2026)
DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged Learning
by: Yoon, Junho, et al.
Published: (2025)
by: Yoon, Junho, et al.
Published: (2025)
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
by: Cho, Janghoon, et al.
Published: (2025)
by: Cho, Janghoon, et al.
Published: (2025)
TokenMotion: Decoupled Motion Control via Token Disentanglement for Human-centric Video Generation
by: Li, Ruineng, et al.
Published: (2025)
by: Li, Ruineng, et al.
Published: (2025)
Stateful Token Reduction for Long-Video Hybrid VLMs
by: Jiang, Jindong, et al.
Published: (2026)
by: Jiang, Jindong, et al.
Published: (2026)
LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
by: Krojer, Benno, et al.
Published: (2026)
by: Krojer, Benno, et al.
Published: (2026)
AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
by: Zhang, Junyang, et al.
Published: (2025)
by: Zhang, Junyang, et al.
Published: (2025)
Similar Items
-
Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization
by: Park, Jungin, et al.
Published: (2025) -
Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations
by: Park, Jungin, et al.
Published: (2025) -
Saliency-Aware Model Merging
by: Park, Jungin, et al.
Published: (2026) -
Bridging Vision and Language Spaces with Assignment Prediction
by: Park, Jungin, et al.
Published: (2024) -
Faster Parameter-Efficient Tuning with Token Redundancy Reduction
by: Kim, Kwonyoung, et al.
Published: (2025)