LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Jihwan, Parthasarathy, Nikhil, Qin, Danfeng, Hur, Junhwa, Sun, Deqing, Han, Bohyung, Yang, Ming-Hsuan, Gong, Boqing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
High-Resolution Frame Interpolation with Patch-based Cascaded Diffusion
di: Hur, Junhwa, et al.
Pubblicazione: (2024)
di: Hur, Junhwa, et al.
Pubblicazione: (2024)
FIFO-Diffusion: Generating Infinite Videos from Text without Training
di: Kim, Jihwan, et al.
Pubblicazione: (2024)
di: Kim, Jihwan, et al.
Pubblicazione: (2024)
Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
di: Chung, Hyungjin, et al.
Pubblicazione: (2025)
di: Chung, Hyungjin, et al.
Pubblicazione: (2025)
Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
di: Kim, Minji, et al.
Pubblicazione: (2025)
di: Kim, Minji, et al.
Pubblicazione: (2025)
Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
di: Zhang, Junyi, et al.
Pubblicazione: (2023)
di: Zhang, Junyi, et al.
Pubblicazione: (2023)
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
di: Zhang, Junyi, et al.
Pubblicazione: (2024)
di: Zhang, Junyi, et al.
Pubblicazione: (2024)
LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
di: Zhang, Junyi, et al.
Pubblicazione: (2026)
di: Zhang, Junyi, et al.
Pubblicazione: (2026)
GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure
di: Gu, Leslie, et al.
Pubblicazione: (2025)
di: Gu, Leslie, et al.
Pubblicazione: (2025)
Restoration-Oriented Video Frame Interpolation with Region-Distinguishable Priors from SAM
di: Han, Yan, et al.
Pubblicazione: (2023)
di: Han, Yan, et al.
Pubblicazione: (2023)
The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
di: Tan, Yuwen, et al.
Pubblicazione: (2025)
di: Tan, Yuwen, et al.
Pubblicazione: (2025)
New Tight Wavelet Frame Constructions Sharing Responsibility
di: Hur, Youngmi, et al.
Pubblicazione: (2024)
di: Hur, Youngmi, et al.
Pubblicazione: (2024)
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
di: Zhang, Shaojie, et al.
Pubblicazione: (2025)
di: Zhang, Shaojie, et al.
Pubblicazione: (2025)
Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
di: Ghazanfari, Sara, et al.
Pubblicazione: (2025)
di: Ghazanfari, Sara, et al.
Pubblicazione: (2025)
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
di: Kim, Namho, et al.
Pubblicazione: (2025)
di: Kim, Namho, et al.
Pubblicazione: (2025)
Communication-Efficient Federated Learning with Accelerated Client Gradient
di: Kim, Geeho, et al.
Pubblicazione: (2022)
di: Kim, Geeho, et al.
Pubblicazione: (2022)
Leveraging Temporal Contextualization for Video Action Recognition
di: Kim, Minji, et al.
Pubblicazione: (2024)
di: Kim, Minji, et al.
Pubblicazione: (2024)
Boundary Attention: Learning curves, corners, junctions and grouping
di: Polansky, Mia Gaia, et al.
Pubblicazione: (2024)
di: Polansky, Mia Gaia, et al.
Pubblicazione: (2024)
GP-4DGS: Probabilistic 4D Gaussian Splatting from Monocular Video via Variational Gaussian Processes
di: Kim, Mijeong, et al.
Pubblicazione: (2026)
di: Kim, Mijeong, et al.
Pubblicazione: (2026)
LADDER: An Efficient Framework for Video Frame Interpolation
di: Shen, Tong, et al.
Pubblicazione: (2024)
di: Shen, Tong, et al.
Pubblicazione: (2024)
Motion-Aware Video Frame Interpolation
di: Han, Pengfei, et al.
Pubblicazione: (2024)
di: Han, Pengfei, et al.
Pubblicazione: (2024)
Emergent Temporal Correspondences from Video Diffusion Transformers
di: Nam, Jisu, et al.
Pubblicazione: (2025)
di: Nam, Jisu, et al.
Pubblicazione: (2025)
UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images
di: Hur, Junhwa, et al.
Pubblicazione: (2026)
di: Hur, Junhwa, et al.
Pubblicazione: (2026)
Re-evaluating Group Robustness via Adaptive Class-Specific Scaling
di: Seo, Seonguk, et al.
Pubblicazione: (2024)
di: Seo, Seonguk, et al.
Pubblicazione: (2024)
Tinted Frames: Question Framing Blinds Vision-Language Models
di: Fan, Wan-Cyuan, et al.
Pubblicazione: (2026)
di: Fan, Wan-Cyuan, et al.
Pubblicazione: (2026)
Frame by Familiar Frame: Understanding Replication in Video Diffusion Models
di: Rahman, Aimon, et al.
Pubblicazione: (2024)
di: Rahman, Aimon, et al.
Pubblicazione: (2024)
Image Diffusion Preview with Consistency Solver
di: Wang, Fu-Yun, et al.
Pubblicazione: (2025)
di: Wang, Fu-Yun, et al.
Pubblicazione: (2025)
Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models
di: Jang, Sangwon, et al.
Pubblicazione: (2025)
di: Jang, Sangwon, et al.
Pubblicazione: (2025)
Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions
di: Hur, Chan, et al.
Pubblicazione: (2025)
di: Hur, Chan, et al.
Pubblicazione: (2025)
Frame Scaling by Graphs
di: K, Ayyanar, et al.
Pubblicazione: (2024)
di: K, Ayyanar, et al.
Pubblicazione: (2024)
Do LLMs Encode Frame Semantics? Evidence from Frame Identification
di: Chundru, Jayanth Krishna, et al.
Pubblicazione: (2025)
di: Chundru, Jayanth Krishna, et al.
Pubblicazione: (2025)
A Simple Recipe for Contrastively Pre-training Video-First Encoders Beyond 16 Frames
di: Papalampidi, Pinelopi, et al.
Pubblicazione: (2023)
di: Papalampidi, Pinelopi, et al.
Pubblicazione: (2023)
TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking
di: Nam, Jisu, et al.
Pubblicazione: (2026)
di: Nam, Jisu, et al.
Pubblicazione: (2026)
Beyond World-Frame Action Heads: Motion-Centric Action Frames for Vision-Language-Action Models
di: Yang, Huoren, et al.
Pubblicazione: (2026)
di: Yang, Huoren, et al.
Pubblicazione: (2026)
Frame by Frame
di: Frank, Hannah
Pubblicazione: (2019)
di: Frank, Hannah
Pubblicazione: (2019)
Frame by Frame
di: Frank, Hannah
Pubblicazione: (2020)
di: Frank, Hannah
Pubblicazione: (2020)
MASIV: Toward Material-Agnostic System Identification from Videos
di: Zhao, Yizhou, et al.
Pubblicazione: (2025)
di: Zhao, Yizhou, et al.
Pubblicazione: (2025)
VideoPrism: A Foundational Visual Encoder for Video Understanding
di: Zhao, Long, et al.
Pubblicazione: (2024)
di: Zhao, Long, et al.
Pubblicazione: (2024)
Frame-Difference Guided Dynamic Region Perception for CLIP Adaptation in Text-Video Retrieval
di: Yu, Jiaao, et al.
Pubblicazione: (2025)
di: Yu, Jiaao, et al.
Pubblicazione: (2025)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
di: Ge, Haonan, et al.
Pubblicazione: (2025)
di: Ge, Haonan, et al.
Pubblicazione: (2025)
Frame-Voyager: Learning to Query Frames for Video Large Language Models
di: Yu, Sicheng, et al.
Pubblicazione: (2024)
di: Yu, Sicheng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
High-Resolution Frame Interpolation with Patch-based Cascaded Diffusion
di: Hur, Junhwa, et al.
Pubblicazione: (2024) -
FIFO-Diffusion: Generating Infinite Videos from Text without Training
di: Kim, Jihwan, et al.
Pubblicazione: (2024) -
Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
di: Chung, Hyungjin, et al.
Pubblicazione: (2025) -
Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
di: Kim, Minji, et al.
Pubblicazione: (2025) -
Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
di: Zhang, Junyi, et al.
Pubblicazione: (2023)