Espresso: High Compression For Rich Extraction From Videos for Your Vision-Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Keunwoo Peter, Dave, Achal, Ambrus, Rares, Mercat, Jean |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GRIN: Zero-Shot Metric Depth with Pixel-Level Diffusion
von: Guizilini, Vitor, et al.
Veröffentlicht: (2024)
von: Guizilini, Vitor, et al.
Veröffentlicht: (2024)
Understanding Video Transformers via Universal Concept Discovery
von: Kowal, Matthew, et al.
Veröffentlicht: (2024)
von: Kowal, Matthew, et al.
Veröffentlicht: (2024)
Understanding Complexity in VideoQA via Visual Program Generation
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2025)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2025)
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models
von: Yu, Keunwoo Peter, et al.
Veröffentlicht: (2025)
von: Yu, Keunwoo Peter, et al.
Veröffentlicht: (2025)
Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties
von: Yu, Keunwoo Peter, et al.
Veröffentlicht: (2023)
von: Yu, Keunwoo Peter, et al.
Veröffentlicht: (2023)
AllTracker: Efficient Dense Point Tracking at High Resolution
von: Harley, Adam W., et al.
Veröffentlicht: (2025)
von: Harley, Adam W., et al.
Veröffentlicht: (2025)
Should VLMs be Pre-trained with Image Data?
von: Keh, Sedrick, et al.
Veröffentlicht: (2025)
von: Keh, Sedrick, et al.
Veröffentlicht: (2025)
Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma?
von: Qu, Tianyuan, et al.
Veröffentlicht: (2025)
von: Qu, Tianyuan, et al.
Veröffentlicht: (2025)
ReFiNe: Recursive Field Networks for Cross-modal Multi-scene Representation
von: Zakharov, Sergey, et al.
Veröffentlicht: (2024)
von: Zakharov, Sergey, et al.
Veröffentlicht: (2024)
Espresso: Robust Concept Filtering in Text-to-Image Models
von: Das, Anudeep, et al.
Veröffentlicht: (2024)
von: Das, Anudeep, et al.
Veröffentlicht: (2024)
EVE: Towards End-to-End Video Subtitle Extraction with Vision-Language Models
von: Yu, Haiyang, et al.
Veröffentlicht: (2025)
von: Yu, Haiyang, et al.
Veröffentlicht: (2025)
Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion
von: Guizilini, Vitor, et al.
Veröffentlicht: (2025)
von: Guizilini, Vitor, et al.
Veröffentlicht: (2025)
Transcrib3D: 3D Referring Expression Resolution through Large Language Models
von: Fang, Jiading, et al.
Veröffentlicht: (2024)
von: Fang, Jiading, et al.
Veröffentlicht: (2024)
$SE(3)$ Equivariant Ray Embeddings for Implicit Multi-View Depth Estimation
von: Xu, Yinshuang, et al.
Veröffentlicht: (2024)
von: Xu, Yinshuang, et al.
Veröffentlicht: (2024)
Zero-Shot Multi-Object Scene Completion
von: Iwase, Shun, et al.
Veröffentlicht: (2024)
von: Iwase, Shun, et al.
Veröffentlicht: (2024)
OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real World
von: Liu, Katherine, et al.
Veröffentlicht: (2025)
von: Liu, Katherine, et al.
Veröffentlicht: (2025)
One Last Attention for Your Vision-Language Model
von: Chen, Liang, et al.
Veröffentlicht: (2025)
von: Chen, Liang, et al.
Veröffentlicht: (2025)
UVG-VPC: Voxelized Point Cloud Dataset for Visual Volumetric Video-based Coding
von: Gautier, Guillaume, et al.
Veröffentlicht: (2025)
von: Gautier, Guillaume, et al.
Veröffentlicht: (2025)
How to Design and Train Your Implicit Neural Representation for Video Compression
von: Gwilliam, Matthew, et al.
Veröffentlicht: (2025)
von: Gwilliam, Matthew, et al.
Veröffentlicht: (2025)
Zero-Shot Open-Vocabulary Tracking with Large Pre-Trained Models
von: Chu, Wen-Hsuan, et al.
Veröffentlicht: (2023)
von: Chu, Wen-Hsuan, et al.
Veröffentlicht: (2023)
Dreamitate: Real-World Visuomotor Policy Learning via Video Generation
von: Liang, Junbang, et al.
Veröffentlicht: (2024)
von: Liang, Junbang, et al.
Veröffentlicht: (2024)
Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks
von: Jia, Mengzhao, et al.
Veröffentlicht: (2024)
von: Jia, Mengzhao, et al.
Veröffentlicht: (2024)
Attention! Your Vision Language Model Could Be Maliciously Manipulated
von: Wang, Xiaosen, et al.
Veröffentlicht: (2025)
von: Wang, Xiaosen, et al.
Veröffentlicht: (2025)
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
von: Han, Songhao, et al.
Veröffentlicht: (2024)
von: Han, Songhao, et al.
Veröffentlicht: (2024)
Enhancing Vision-Language Pre-training with Rich Supervisions
von: Gao, Yuan, et al.
Veröffentlicht: (2024)
von: Gao, Yuan, et al.
Veröffentlicht: (2024)
From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos
von: Gupta, Animesh, et al.
Veröffentlicht: (2025)
von: Gupta, Animesh, et al.
Veröffentlicht: (2025)
Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
DiffusionNOCS: Managing Symmetry and Uncertainty in Sim2Real Multi-Modal Category-level Pose Estimation
von: Ikeda, Takuya, et al.
Veröffentlicht: (2024)
von: Ikeda, Takuya, et al.
Veröffentlicht: (2024)
Language-Guided Token Compression with Reinforcement Learning in Large Vision-Language Models
von: Cao, Sihan, et al.
Veröffentlicht: (2026)
von: Cao, Sihan, et al.
Veröffentlicht: (2026)
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
Distilling Vision-Language Models on Millions of Videos
von: Zhao, Yue, et al.
Veröffentlicht: (2024)
von: Zhao, Yue, et al.
Veröffentlicht: (2024)
Hierarchical Reasoning with Vision-Language Models for Incident Reports from Dashcam Videos
von: Yokoi, Shingo, et al.
Veröffentlicht: (2025)
von: Yokoi, Shingo, et al.
Veröffentlicht: (2025)
Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
Extreme Model Compression for Edge Vision-Language Models: Sparse Temporal Token Fusion and Adaptive Neural Compression
von: Tanvir, Md Tasnin, et al.
Veröffentlicht: (2025)
von: Tanvir, Md Tasnin, et al.
Veröffentlicht: (2025)
Vision-centric Token Compression in Large Language Model
von: Xing, Ling, et al.
Veröffentlicht: (2025)
von: Xing, Ling, et al.
Veröffentlicht: (2025)
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
Long-Tailed 3D Detection via Multi-Modal Fusion
von: Ma, Yechi, et al.
Veröffentlicht: (2023)
von: Ma, Yechi, et al.
Veröffentlicht: (2023)
NeRF-MAE: Masked AutoEncoders for Self-Supervised 3D Representation Learning for Neural Radiance Fields
von: Irshad, Muhammad Zubair, et al.
Veröffentlicht: (2024)
von: Irshad, Muhammad Zubair, et al.
Veröffentlicht: (2024)
Inference Compute-Optimal Video Vision Language Models
von: Wang, Peiqi, et al.
Veröffentlicht: (2025)
von: Wang, Peiqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GRIN: Zero-Shot Metric Depth with Pixel-Level Diffusion
von: Guizilini, Vitor, et al.
Veröffentlicht: (2024) -
Understanding Video Transformers via Universal Concept Discovery
von: Kowal, Matthew, et al.
Veröffentlicht: (2024) -
Understanding Complexity in VideoQA via Visual Program Generation
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2025) -
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models
von: Yu, Keunwoo Peter, et al.
Veröffentlicht: (2025) -
Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties
von: Yu, Keunwoo Peter, et al.
Veröffentlicht: (2023)