SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Biyyala, Varun, Kathuria, Bharat Chanderprakash, Li, Jialu, Zhang, Youshan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SparrowVQE: Visual Question Explanation for Course Content Understanding
by: Li, Jialu, et al.
Published: (2024)
by: Li, Jialu, et al.
Published: (2024)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
by: Cheng, Zesen, et al.
Published: (2024)
by: Cheng, Zesen, et al.
Published: (2024)
Automatic Teaching Platform on Vision Language Retrieval Augmented Generation
by: Gokhman, Ruslan, et al.
Published: (2025)
by: Gokhman, Ruslan, et al.
Published: (2025)
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models
by: Liang, Yiming, et al.
Published: (2026)
by: Liang, Yiming, et al.
Published: (2026)
FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition
by: Yu, Enhui, et al.
Published: (2026)
by: Yu, Enhui, et al.
Published: (2026)
EVQAScore: A Fine-grained Metric for Video Question Answering Data Quality Evaluation
by: Liang, Hao, et al.
Published: (2024)
by: Liang, Hao, et al.
Published: (2024)
STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question Answering
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation
by: Li, Jialu, et al.
Published: (2024)
by: Li, Jialu, et al.
Published: (2024)
Temporal Reasoning Transfer from Text to Video
by: Li, Lei, et al.
Published: (2024)
by: Li, Lei, et al.
Published: (2024)
Spatial Semantic Recurrent Mining for Referring Image Segmentation
by: Yang, Jiaxing, et al.
Published: (2024)
by: Yang, Jiaxing, et al.
Published: (2024)
Error-Driven Scene Editing for 3D Grounding in Large Language Models
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
LLMs Meet Long Video: Advancing Long Video Question Answering with An Interactive Visual Adapter in LLMs
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis
by: Dipta, Shubhashis Roy, et al.
Published: (2025)
by: Dipta, Shubhashis Roy, et al.
Published: (2025)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
by: Kasaei, Seyed Amir, et al.
Published: (2025)
by: Kasaei, Seyed Amir, et al.
Published: (2025)
SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
by: Kendre, Shrikant, et al.
Published: (2025)
by: Kendre, Shrikant, et al.
Published: (2025)
From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding
by: Wang, Xiangfeng, et al.
Published: (2025)
by: Wang, Xiangfeng, et al.
Published: (2025)
Unhackable Temporal Rewarding for Scalable Video MLLMs
by: Yu, En, et al.
Published: (2025)
by: Yu, En, et al.
Published: (2025)
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
S3Editor: A Sparse Semantic-Disentangled Self-Training Framework for Face Video Editing
by: Wang, Guangzhi, et al.
Published: (2024)
by: Wang, Guangzhi, et al.
Published: (2024)
Attribution Analysis Meets Model Editing: Advancing Knowledge Correction in Vision Language Models with VisEdit
by: Chen, Qizhou, et al.
Published: (2024)
by: Chen, Qizhou, et al.
Published: (2024)
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
by: Li, Yun, et al.
Published: (2025)
by: Li, Yun, et al.
Published: (2025)
Zero-Shot Video Translation and Editing with Frame Spatial-Temporal Correspondence
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
by: Hu, Kairui, et al.
Published: (2025)
by: Hu, Kairui, et al.
Published: (2025)
A Video-grounded Dialogue Dataset and Metric for Event-driven Activities
by: Imrattanatrai, Wiradee, et al.
Published: (2025)
by: Imrattanatrai, Wiradee, et al.
Published: (2025)
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
by: Luo, Fuwen, et al.
Published: (2025)
by: Luo, Fuwen, et al.
Published: (2025)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
by: Wang, Zun, et al.
Published: (2024)
by: Wang, Zun, et al.
Published: (2024)
FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion Editing
by: Premsri, Tanawan, et al.
Published: (2025)
by: Premsri, Tanawan, et al.
Published: (2025)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
by: Liu, Xuannan, et al.
Published: (2025)
by: Liu, Xuannan, et al.
Published: (2025)
A Video is Worth 256 Bases: Spatial-Temporal Expectation-Maximization Inversion for Zero-Shot Video Editing
by: Li, Maomao, et al.
Published: (2023)
by: Li, Maomao, et al.
Published: (2023)
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
by: Fu, Chaoyou, et al.
Published: (2024)
by: Fu, Chaoyou, et al.
Published: (2024)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
by: Li, Sifan, et al.
Published: (2025)
by: Li, Sifan, et al.
Published: (2025)
Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis
by: Liu, Runzhou, et al.
Published: (2026)
by: Liu, Runzhou, et al.
Published: (2026)
Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times
by: Loginova, Olga, et al.
Published: (2025)
by: Loginova, Olga, et al.
Published: (2025)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
by: Li, Jialu, et al.
Published: (2025)
by: Li, Jialu, et al.
Published: (2025)
Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models
by: Lu, Jiaying, et al.
Published: (2023)
by: Lu, Jiaying, et al.
Published: (2023)
Tracing the Arrow of Time: Diagnosing Temporal Information Flow in Video-LLMs
by: Han, Peitao, et al.
Published: (2026)
by: Han, Peitao, et al.
Published: (2026)
Xiaoice: Training-Free Video Understanding via Self-Supervised Spatio-Temporal Clustering of Semantic Features
by: Ji, Shihao, et al.
Published: (2025)
by: Ji, Shihao, et al.
Published: (2025)
Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning
by: Wei, Xinyu, et al.
Published: (2025)
by: Wei, Xinyu, et al.
Published: (2025)
Semantic Frame Aggregation-based Transformer for Live Video Comment Generation
by: Fatima, Anam, et al.
Published: (2025)
by: Fatima, Anam, et al.
Published: (2025)
Similar Items
-
SparrowVQE: Visual Question Explanation for Course Content Understanding
by: Li, Jialu, et al.
Published: (2024) -
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
by: Cheng, Zesen, et al.
Published: (2024) -
Automatic Teaching Platform on Vision Language Retrieval Augmented Generation
by: Gokhman, Ruslan, et al.
Published: (2025) -
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models
by: Liang, Yiming, et al.
Published: (2026) -
FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition
by: Yu, Enhui, et al.
Published: (2026)