Xiaoice: Training-Free Video Understanding via Self-Supervised Spatio-Temporal Clustering of Semantic Features
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ji, Shihao, Song, Zihui |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
LFTR: Learning-Free Token Reduction for Multimodal Large Language Models
par: Zhao, Zihui, et autres
Publié: (2025)
par: Zhao, Zihui, et autres
Publié: (2025)
S3Editor: A Sparse Semantic-Disentangled Self-Training Framework for Face Video Editing
par: Wang, Guangzhi, et autres
Publié: (2024)
par: Wang, Guangzhi, et autres
Publié: (2024)
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
par: Hyun, Jeongseok, et autres
Publié: (2025)
par: Hyun, Jeongseok, et autres
Publié: (2025)
STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing
par: Lee, Junsung, et autres
Publié: (2025)
par: Lee, Junsung, et autres
Publié: (2025)
Spatio-Temporal Side Tuning Pre-trained Foundation Models for Video-based Pedestrian Attribute Recognition
par: Wang, Xiao, et autres
Publié: (2024)
par: Wang, Xiao, et autres
Publié: (2024)
Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly
par: Chetan, Aditya, et autres
Publié: (2026)
par: Chetan, Aditya, et autres
Publié: (2026)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
par: Fu, Honghao, et autres
Publié: (2026)
par: Fu, Honghao, et autres
Publié: (2026)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
par: Yuan, Huaying, et autres
Publié: (2025)
par: Yuan, Huaying, et autres
Publié: (2025)
VITED: Video Temporal Evidence Distillation
par: Lu, Yujie, et autres
Publié: (2025)
par: Lu, Yujie, et autres
Publié: (2025)
Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal Prompts
par: Wu, Peng, et autres
Publié: (2024)
par: Wu, Peng, et autres
Publié: (2024)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
par: Lee, Daeun, et autres
Publié: (2025)
par: Lee, Daeun, et autres
Publié: (2025)
VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
par: Li, Shicheng, et autres
Publié: (2023)
par: Li, Shicheng, et autres
Publié: (2023)
FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding
par: Cao, Zhuo, et autres
Publié: (2024)
par: Cao, Zhuo, et autres
Publié: (2024)
Adaptive Greedy Frame Selection for Long Video Understanding
par: Huang, Yuning, et autres
Publié: (2026)
par: Huang, Yuning, et autres
Publié: (2026)
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
par: Wang, Ye, et autres
Publié: (2025)
par: Wang, Ye, et autres
Publié: (2025)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
par: Wang, Dianyi, et autres
Publié: (2025)
par: Wang, Dianyi, et autres
Publié: (2025)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
par: Wang, Shihao, et autres
Publié: (2025)
par: Wang, Shihao, et autres
Publié: (2025)
SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards
par: Hong, Jixiang, et autres
Publié: (2025)
par: Hong, Jixiang, et autres
Publié: (2025)
PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection
par: Bi, Jinhe, et autres
Publié: (2025)
par: Bi, Jinhe, et autres
Publié: (2025)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
par: Xu, Ruyi, et autres
Publié: (2025)
par: Xu, Ruyi, et autres
Publié: (2025)
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
par: Cai, Mu, et autres
Publié: (2024)
par: Cai, Mu, et autres
Publié: (2024)
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation
par: Xie, Yichen, et autres
Publié: (2025)
par: Xie, Yichen, et autres
Publié: (2025)
Clustering via Self-Supervised Diffusion
par: Uziel, Roy, et autres
Publié: (2025)
par: Uziel, Roy, et autres
Publié: (2025)
VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction
par: Wang, Hao, et autres
Publié: (2025)
par: Wang, Hao, et autres
Publié: (2025)
Semantic Matters: Multimodal Features for Affective Analysis
par: Hallmen, Tobias, et autres
Publié: (2025)
par: Hallmen, Tobias, et autres
Publié: (2025)
Self-Supervised Learning Based Handwriting Verification
par: Chauhan, Mihir, et autres
Publié: (2024)
par: Chauhan, Mihir, et autres
Publié: (2024)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
par: Li, Yunxin, et autres
Publié: (2024)
par: Li, Yunxin, et autres
Publié: (2024)
FILS: Self-Supervised Video Feature Prediction In Semantic Language Space
par: Ahmadian, Mona, et autres
Publié: (2024)
par: Ahmadian, Mona, et autres
Publié: (2024)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
par: Feng, Weixi, et autres
Publié: (2024)
par: Feng, Weixi, et autres
Publié: (2024)
MMVU: Measuring Expert-Level Multi-Discipline Video Understanding
par: Zhao, Yilun, et autres
Publié: (2025)
par: Zhao, Yilun, et autres
Publié: (2025)
Learning Spatio-Temporal Feature Representations for Video-Based Gaze Estimation
par: Personnic, Alexandre, et autres
Publié: (2025)
par: Personnic, Alexandre, et autres
Publié: (2025)
APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
par: Huang, Yuxiang, et autres
Publié: (2026)
par: Huang, Yuxiang, et autres
Publié: (2026)
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment
par: Li, Wei, et autres
Publié: (2024)
par: Li, Wei, et autres
Publié: (2024)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
par: Wang, Ziyang, et autres
Publié: (2025)
par: Wang, Ziyang, et autres
Publié: (2025)
HawkEye: Training Video-Text LLMs for Grounding Text in Videos
par: Wang, Yueqian, et autres
Publié: (2024)
par: Wang, Yueqian, et autres
Publié: (2024)
SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding
par: Wu, Chang-Hsun, et autres
Publié: (2025)
par: Wu, Chang-Hsun, et autres
Publié: (2025)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
par: Li, Jialu, et autres
Publié: (2025)
par: Li, Jialu, et autres
Publié: (2025)
Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
par: Lee, Dohun, et autres
Publié: (2025)
par: Lee, Dohun, et autres
Publié: (2025)
Temporal Preference Optimization for Long-Form Video Understanding
par: Li, Rui, et autres
Publié: (2025)
par: Li, Rui, et autres
Publié: (2025)
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
par: Zhang, Xiaoyi, et autres
Publié: (2025)
par: Zhang, Xiaoyi, et autres
Publié: (2025)
Documents similaires
-
LFTR: Learning-Free Token Reduction for Multimodal Large Language Models
par: Zhao, Zihui, et autres
Publié: (2025) -
S3Editor: A Sparse Semantic-Disentangled Self-Training Framework for Face Video Editing
par: Wang, Guangzhi, et autres
Publié: (2024) -
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
par: Hyun, Jeongseok, et autres
Publié: (2025) -
STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing
par: Lee, Junsung, et autres
Publié: (2025) -
Spatio-Temporal Side Tuning Pre-trained Foundation Models for Video-based Pedestrian Attribute Recognition
par: Wang, Xiao, et autres
Publié: (2024)