FreeVA: Offline MLLM as Training-Free Video Assistant
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Wu, Wenhao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound
von: Guo, Xuechen, et al.
Veröffentlicht: (2024)
von: Guo, Xuechen, et al.
Veröffentlicht: (2024)
A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images
von: Lee, Jaeseong, et al.
Veröffentlicht: (2025)
von: Lee, Jaeseong, et al.
Veröffentlicht: (2025)
EvoStreaming: Your Offline Video Model Is a Natively Streaming Assistant
von: Wen, Zichen, et al.
Veröffentlicht: (2026)
von: Wen, Zichen, et al.
Veröffentlicht: (2026)
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
von: Sun, Bowen, et al.
Veröffentlicht: (2025)
von: Sun, Bowen, et al.
Veröffentlicht: (2025)
Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs
von: Han, Kai, et al.
Veröffentlicht: (2024)
von: Han, Kai, et al.
Veröffentlicht: (2024)
LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning
von: Li, Jiajie, et al.
Veröffentlicht: (2024)
von: Li, Jiajie, et al.
Veröffentlicht: (2024)
Geometry-Aware Semantic Reasoning for Training Free Video Anomaly Detection
von: Zia, Ali, et al.
Veröffentlicht: (2026)
von: Zia, Ali, et al.
Veröffentlicht: (2026)
Block Cascading: Training Free Acceleration of Block-Causal Video Models
von: Bandyopadhyay, Hmrishav, et al.
Veröffentlicht: (2025)
von: Bandyopadhyay, Hmrishav, et al.
Veröffentlicht: (2025)
Vocabulary-Free 3D Instance Segmentation with Vision and Language Assistant
von: Mei, Guofeng, et al.
Veröffentlicht: (2024)
von: Mei, Guofeng, et al.
Veröffentlicht: (2024)
FreeSliders: Training-Free, Modality-Agnostic Concept Sliders for Fine-Grained Diffusion Control in Images, Audio, and Video
von: Ezra, Rotem, et al.
Veröffentlicht: (2025)
von: Ezra, Rotem, et al.
Veröffentlicht: (2025)
Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation
von: Zhao, Bingrui, et al.
Veröffentlicht: (2025)
von: Zhao, Bingrui, et al.
Veröffentlicht: (2025)
DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM
von: Wu, Yixuan, et al.
Veröffentlicht: (2024)
von: Wu, Yixuan, et al.
Veröffentlicht: (2024)
FAIRT2V: Training-Free Debiasing for Text-to-Video Diffusion Models
von: Zhong, Haonan, et al.
Veröffentlicht: (2026)
von: Zhong, Haonan, et al.
Veröffentlicht: (2026)
Ada-VE: Training-Free Consistent Video Editing Using Adaptive Motion Prior
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
von: Hyun, Jeongseok, et al.
Veröffentlicht: (2025)
von: Hyun, Jeongseok, et al.
Veröffentlicht: (2025)
STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing
von: Lee, Junsung, et al.
Veröffentlicht: (2025)
von: Lee, Junsung, et al.
Veröffentlicht: (2025)
Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models
von: Jang, Sangwon, et al.
Veröffentlicht: (2025)
von: Jang, Sangwon, et al.
Veröffentlicht: (2025)
FreeOrbit4D: Training-Free Arbitrary Camera Redirection for Monocular Videos via Foreground-Complete 4D Reconstruction
von: Cao, Wei, et al.
Veröffentlicht: (2026)
von: Cao, Wei, et al.
Veröffentlicht: (2026)
StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
von: Wang, Haibo, et al.
Veröffentlicht: (2025)
von: Wang, Haibo, et al.
Veröffentlicht: (2025)
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
von: Tzachor, Issar, et al.
Veröffentlicht: (2026)
von: Tzachor, Issar, et al.
Veröffentlicht: (2026)
Free$^2$Guide: Training-Free Text-to-Video Alignment using Image LVLM
von: Kim, Jaemin, et al.
Veröffentlicht: (2024)
von: Kim, Jaemin, et al.
Veröffentlicht: (2024)
Training-Free Unsupervised Prompt for Vision-Language Models
von: Long, Sifan, et al.
Veröffentlicht: (2024)
von: Long, Sifan, et al.
Veröffentlicht: (2024)
LIME: Less Is More for MLLM Evaluation
von: Zhu, King, et al.
Veröffentlicht: (2024)
von: Zhu, King, et al.
Veröffentlicht: (2024)
RASA: Replace Anyone, Say Anything -- A Training-Free Framework for Audio-Driven and Universal Portrait Video Editing
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
EchoPilot: Training-Free Ultrasound Video Segmentation via Scale-Space Semantic Prompting and Reliability-Gated Memory
von: Xiao, Ruiqiang, et al.
Veröffentlicht: (2026)
von: Xiao, Ruiqiang, et al.
Veröffentlicht: (2026)
HASTE: Training-Free Video Diffusion Acceleration via Head-Wise Adaptive Sparse Attention
von: Zheng, Xuzhe, et al.
Veröffentlicht: (2026)
von: Zheng, Xuzhe, et al.
Veröffentlicht: (2026)
D3: Training-Free AI-Generated Video Detection Using Second-Order Features
von: Zheng, Chende, et al.
Veröffentlicht: (2025)
von: Zheng, Chende, et al.
Veröffentlicht: (2025)
CoReVAD: A Contextual Reasoning Framework for Training-Free Video Anomaly Detection
von: Lim, Hyeongmuk, et al.
Veröffentlicht: (2026)
von: Lim, Hyeongmuk, et al.
Veröffentlicht: (2026)
BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models
von: Shi, Fengyuan, et al.
Veröffentlicht: (2023)
von: Shi, Fengyuan, et al.
Veröffentlicht: (2023)
EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation
von: Vandersanden, Jente, et al.
Veröffentlicht: (2026)
von: Vandersanden, Jente, et al.
Veröffentlicht: (2026)
VISTA: Mitigating Semantic Inertia in Video-LLMs via Training-Free Dynamic Chain-of-Thought Routing
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
Coupled Diffusion Sampling for Training-Free Multi-View Image Editing
von: Alzayer, Hadi, et al.
Veröffentlicht: (2025)
von: Alzayer, Hadi, et al.
Veröffentlicht: (2025)
GeoSeg: Training-Free Reasoning-Driven Segmentation in Remote Sensing Imagery
von: Jiang, Lifan, et al.
Veröffentlicht: (2026)
von: Jiang, Lifan, et al.
Veröffentlicht: (2026)
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
von: Jin, Yizhang, et al.
Veröffentlicht: (2024)
von: Jin, Yizhang, et al.
Veröffentlicht: (2024)
VGGT-CD: Training-Free Robust Registration for 3D Change Detection
von: Zhang, Wei, et al.
Veröffentlicht: (2026)
von: Zhang, Wei, et al.
Veröffentlicht: (2026)
V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents
von: Yue, Zhengrong, et al.
Veröffentlicht: (2025)
von: Yue, Zhengrong, et al.
Veröffentlicht: (2025)
Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
von: Gu, Bohai, et al.
Veröffentlicht: (2026)
von: Gu, Bohai, et al.
Veröffentlicht: (2026)
TAMEing Long Contexts in Personalization: Towards Training-Free and State-Aware MLLM Personalized Assistant
von: Hong, Rongpei, et al.
Veröffentlicht: (2025)
von: Hong, Rongpei, et al.
Veröffentlicht: (2025)
BAMI: Training-Free Bias Mitigation in GUI Grounding
von: Zhang, Borui, et al.
Veröffentlicht: (2026)
von: Zhang, Borui, et al.
Veröffentlicht: (2026)
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2024)
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound
von: Guo, Xuechen, et al.
Veröffentlicht: (2024) -
A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images
von: Lee, Jaeseong, et al.
Veröffentlicht: (2025) -
EvoStreaming: Your Offline Video Model Is a Natively Streaming Assistant
von: Wen, Zichen, et al.
Veröffentlicht: (2026) -
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
von: Sun, Bowen, et al.
Veröffentlicht: (2025) -
Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs
von: Han, Kai, et al.
Veröffentlicht: (2024)