RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Hang, Cai, Yujun, Ge, Haonan, Chen, Hongkai, Yang, Ming-Hsuan, Wang, Yiwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Structured Attention Matters to Multimodal LLMs in Document Understanding
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
by: Ge, Haonan, et al.
Published: (2025)
by: Ge, Haonan, et al.
Published: (2025)
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
by: Wu, Hang, et al.
Published: (2026)
by: Wu, Hang, et al.
Published: (2026)
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
by: Wu, Hang, et al.
Published: (2025)
by: Wu, Hang, et al.
Published: (2025)
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning
by: Wu, Hang, et al.
Published: (2026)
by: Wu, Hang, et al.
Published: (2026)
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
by: Sun, Bowen, et al.
Published: (2025)
by: Sun, Bowen, et al.
Published: (2025)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
by: Ge, Haonan, et al.
Published: (2025)
by: Ge, Haonan, et al.
Published: (2025)
AudioRouter: Data Efficient Audio Understanding via RL based Dual Reasoning
by: Chen, Liyang, et al.
Published: (2026)
by: Chen, Liyang, et al.
Published: (2026)
Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models
by: Wang, Zhaochen, et al.
Published: (2025)
by: Wang, Zhaochen, et al.
Published: (2025)
AudioMotionBench: Evaluating Auditory Motion Perception in Audio LLMs
by: Sun, Zhe, et al.
Published: (2025)
by: Sun, Zhe, et al.
Published: (2025)
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
by: Wu, Yike, et al.
Published: (2025)
by: Wu, Yike, et al.
Published: (2025)
BRIGHT+: Upgrading the BRIGHT Benchmark with MARCUS, a Multi-Agent RAG Clean-Up Suite
by: Chen, Liyang, et al.
Published: (2025)
by: Chen, Liyang, et al.
Published: (2025)
Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression
by: Cui, Yu, et al.
Published: (2025)
by: Cui, Yu, et al.
Published: (2025)
Cure or Poison? Embedding Instructions Visually Alters Hallucination in Vision-Language Models
by: Wang, Zhaochen, et al.
Published: (2025)
by: Wang, Zhaochen, et al.
Published: (2025)
Finding Distributed Object-Centric Properties in Self-Supervised Transformers
by: Rawlekar, Samyak, et al.
Published: (2026)
by: Rawlekar, Samyak, et al.
Published: (2026)
Process or Result? Manipulated Ending Tokens Can Mislead Reasoning LLMs to Ignore the Correct Reasoning Steps
by: Cui, Yu, et al.
Published: (2025)
by: Cui, Yu, et al.
Published: (2025)
PromptCD: Test-Time Behavior Enhancement via Polarity-Prompt Contrastive Decoding
by: Bi, Baolong, et al.
Published: (2026)
by: Bi, Baolong, et al.
Published: (2026)
Primacy Effect of ChatGPT
by: Wang, Yiwei, et al.
Published: (2023)
by: Wang, Yiwei, et al.
Published: (2023)
Manual2Skill: Learning to Read Manuals and Acquire Robotic Skills for Furniture Assembly Using Vision-Language Models
by: Tie, Chenrui, et al.
Published: (2025)
by: Tie, Chenrui, et al.
Published: (2025)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
by: Fu, Honghao, et al.
Published: (2026)
by: Fu, Honghao, et al.
Published: (2026)
Geological Everything Model 3D: A Promptable Foundation Model for Unified and Zero-Shot Subsurface Understanding
by: Dou, Yimin, et al.
Published: (2025)
by: Dou, Yimin, et al.
Published: (2025)
Skill Path: Unveiling Language Skills from Circuit Graphs
by: Chen, Hang, et al.
Published: (2024)
by: Chen, Hang, et al.
Published: (2024)
WavefrontDiffusion: Dynamic Decoding Schedule for Improved Reasoning
by: Yang, Haojin, et al.
Published: (2025)
by: Yang, Haojin, et al.
Published: (2025)
Lost in Edits? A $λ$-Compass for AIGC Provenance
by: You, Wenhao, et al.
Published: (2025)
by: You, Wenhao, et al.
Published: (2025)
ContextNav: Towards Agentic Multimodal In-Context Learning
by: Fu, Honghao, et al.
Published: (2025)
by: Fu, Honghao, et al.
Published: (2025)
Self-Manager: Parallel Agent Loop for Long-form Deep Research
by: Xu, Yilong, et al.
Published: (2026)
by: Xu, Yilong, et al.
Published: (2026)
Skill-Critic: Refining Learned Skills for Hierarchical Reinforcement Learning
by: Hao, Ce, et al.
Published: (2023)
by: Hao, Ce, et al.
Published: (2023)
Semantic Segmentation Refiner for Ultrasound Applications with Zero-Shot Foundation Models
by: Indelman, Hedda Cohen, et al.
Published: (2024)
by: Indelman, Hedda Cohen, et al.
Published: (2024)
Guiding Skill Discovery with Foundation Models
by: Yang, Zhao, et al.
Published: (2025)
by: Yang, Zhao, et al.
Published: (2025)
Enhancing LLM Character-Level Manipulation via Divide and Conquer
by: Xiong, Zhen, et al.
Published: (2025)
by: Xiong, Zhen, et al.
Published: (2025)
Rethinking Agent Design: From Top-Down Workflows to Bottom-Up Skill Evolution
by: Du, Jiawei, et al.
Published: (2025)
by: Du, Jiawei, et al.
Published: (2025)
Rethinking Early Stopping: Refine, Then Calibrate
by: Berta, Eugène, et al.
Published: (2025)
by: Berta, Eugène, et al.
Published: (2025)
Understanding Gradient Boosting Classifier: Training, Prediction, and the Role of $γ_j$
by: Chen, Hung-Hsuan
Published: (2024)
by: Chen, Hung-Hsuan
Published: (2024)
Evaluating Multi-Turn Bargain Skills in LLM-Based Seller Agent
by: Wang, Issue Yishu, et al.
Published: (2025)
by: Wang, Issue Yishu, et al.
Published: (2025)
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
by: Wang, Liang, et al.
Published: (2026)
by: Wang, Liang, et al.
Published: (2026)
TimeOmni-VL: Unified Models for Time Series Understanding and Generation
by: Guan, Tong, et al.
Published: (2026)
by: Guan, Tong, et al.
Published: (2026)
Structured Outputs Enable General-Purpose LLMs to be Medical Experts
by: Guo, Guangfu, et al.
Published: (2025)
by: Guo, Guangfu, et al.
Published: (2025)
A Survey of Foundation Models for Music Understanding
by: Li, Wenjun, et al.
Published: (2024)
by: Li, Wenjun, et al.
Published: (2024)
CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models
by: Li, Zhong-Zhi, et al.
Published: (2024)
by: Li, Zhong-Zhi, et al.
Published: (2024)
Rethinking Memory Mechanisms of Foundation Agents in the Second Half: A Survey
by: Huang, Wei-Chieh, et al.
Published: (2026)
by: Huang, Wei-Chieh, et al.
Published: (2026)
Similar Items
-
Structured Attention Matters to Multimodal LLMs in Document Understanding
by: Liu, Chang, et al.
Published: (2025) -
MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
by: Ge, Haonan, et al.
Published: (2025) -
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
by: Wu, Hang, et al.
Published: (2026) -
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
by: Wu, Hang, et al.
Published: (2025) -
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning
by: Wu, Hang, et al.
Published: (2026)