Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Weiyu, Chen, Ziyang, Wang, Shaoguang, He, Jianxiang, Xu, Yijie, Ye, Jinhui, Sun, Ying, Xiong, Hui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
by: Wang, Shaoguang, et al.
Published: (2026)
by: Wang, Shaoguang, et al.
Published: (2026)
Event Camera Demosaicing via Swin Transformer and Pixel-focus Loss
by: Lu, Yunfan, et al.
Published: (2024)
by: Lu, Yunfan, et al.
Published: (2024)
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
by: He, Jianxiang, et al.
Published: (2025)
by: He, Jianxiang, et al.
Published: (2025)
DrawVideo: Generating Long Video from Storyboard Keyframe Sketches
by: Xu, Chuanzhi, et al.
Published: (2026)
by: Xu, Chuanzhi, et al.
Published: (2026)
Less is More: Token-Efficient Video-QA via Adaptive Frame-Pruning and Semantic Graph Integration
by: Wang, Shaoguang, et al.
Published: (2025)
by: Wang, Shaoguang, et al.
Published: (2025)
Leveraging Compressed Frame Sizes For Ultra-Fast Video Classification
by: Han, Yuxing, et al.
Published: (2024)
by: Han, Yuxing, et al.
Published: (2024)
A New Logic For Pediatric Brain Tumor Segmentation
by: Bengtsson, Max, et al.
Published: (2024)
by: Bengtsson, Max, et al.
Published: (2024)
Cost-Efficient Multi-Scale Fovea for Semantic-Based Visual Search Attention
by: Luzio, João, et al.
Published: (2026)
by: Luzio, João, et al.
Published: (2026)
BVI-VFI: A Video Quality Database for Video Frame Interpolation
by: Danier, Duolikun, et al.
Published: (2022)
by: Danier, Duolikun, et al.
Published: (2022)
FCA2: Frame Compression-Aware Autoencoder for Modular and Fast Compressed Video Super-Resolution
by: Wang, Zhaoyang, et al.
Published: (2025)
by: Wang, Zhaoyang, et al.
Published: (2025)
LDMVFI: Video Frame Interpolation with Latent Diffusion Models
by: Danier, Duolikun, et al.
Published: (2023)
by: Danier, Duolikun, et al.
Published: (2023)
A Subjective Quality Study for Video Frame Interpolation
by: Danier, Duolikun, et al.
Published: (2022)
by: Danier, Duolikun, et al.
Published: (2022)
Weak-Mamba-UNet: Visual Mamba Makes CNN and ViT Work Better for Scribble-based Medical Image Segmentation
by: Wang, Ziyang, et al.
Published: (2024)
by: Wang, Ziyang, et al.
Published: (2024)
RGB-Event ISP: The Dataset and Benchmark
by: Lu, Yunfan, et al.
Published: (2025)
by: Lu, Yunfan, et al.
Published: (2025)
Compressed Depth Map Super-Resolution and Restoration: AIM 2024 Challenge Results
by: Conde, Marcos V., et al.
Published: (2024)
by: Conde, Marcos V., et al.
Published: (2024)
TriPSS: A Tri-Modal Keyframe Extraction Framework Using Perceptual, Structural, and Semantic Representations
by: Cakmak, Mert Can, et al.
Published: (2025)
by: Cakmak, Mert Can, et al.
Published: (2025)
Semi-Mamba-UNet: Pixel-Level Contrastive and Pixel-Level Cross-Supervised Visual Mamba-based UNet for Semi-Supervised Medical Image Segmentation
by: Ma, Chao, et al.
Published: (2024)
by: Ma, Chao, et al.
Published: (2024)
Self-Supervised One-Step Diffusion Refinement for Snapshot Compressive Imaging
by: Huang, Shaoguang, et al.
Published: (2024)
by: Huang, Shaoguang, et al.
Published: (2024)
Compressing Human Body Video with Interactive Semantics: A Generative Approach
by: Chen, Bolin, et al.
Published: (2025)
by: Chen, Bolin, et al.
Published: (2025)
Joint Reference Frame Synthesis and Post Filter Enhancement for Versatile Video Coding
by: Bao, Weijie, et al.
Published: (2024)
by: Bao, Weijie, et al.
Published: (2024)
On the Benefits of Visual Stabilization for Frame- and Event-based Perception
by: Rodriguez-Gomez, Juan Pablo, et al.
Published: (2024)
by: Rodriguez-Gomez, Juan Pablo, et al.
Published: (2024)
Neural-Network-Enhanced Metalens Camera for High-Definition, Dynamic Imaging in the Long-Wave Infrared Spectrum
by: Wei, Jing-Yang, et al.
Published: (2024)
by: Wei, Jing-Yang, et al.
Published: (2024)
When CNN Meet with ViT: Towards Semi-Supervised Learning for Multi-Class Medical Image Semantic Segmentation
by: Wang, Ziyang, et al.
Published: (2022)
by: Wang, Ziyang, et al.
Published: (2022)
Semi-Unsupervised Microscopy Segmentation with Fuzzy Logic and Spatial Statistics for Cross-Domain Analysis Using a GUI
by: Das, Surajit, et al.
Published: (2025)
by: Das, Surajit, et al.
Published: (2025)
Fine-Grained Motion Compression and Selective Temporal Fusion for Neural B-Frame Video Coding
by: Sheng, Xihua, et al.
Published: (2025)
by: Sheng, Xihua, et al.
Published: (2025)
Enhancing Deformable Convolution based Video Frame Interpolation with Coarse-to-fine 3D CNN
by: Danier, Duolikun, et al.
Published: (2022)
by: Danier, Duolikun, et al.
Published: (2022)
Fuzzy Logic-Based System for Brain Tumour Detection and Classification
by: Narasimham, NVSL, et al.
Published: (2024)
by: Narasimham, NVSL, et al.
Published: (2024)
Beyond GFVC: A Progressive Face Video Compression Framework with Adaptive Visual Tokens
by: Chen, Bolin, et al.
Published: (2024)
by: Chen, Bolin, et al.
Published: (2024)
SAR-RAG: ATR Visual Question Answering by Semantic Search, Retrieval, and MLLM Generation
by: Ramirez, David F., et al.
Published: (2026)
by: Ramirez, David F., et al.
Published: (2026)
Mamba-UNet: UNet-Like Pure Visual Mamba for Medical Image Segmentation
by: Wang, Ziyang, et al.
Published: (2024)
by: Wang, Ziyang, et al.
Published: (2024)
Turb-Seg-Res: A Segment-then-Restore Pipeline for Dynamic Videos with Atmospheric Turbulence
by: Saha, Ripon Kumar, et al.
Published: (2024)
by: Saha, Ripon Kumar, et al.
Published: (2024)
Deep Diversity-Enhanced Feature Representation of Hyperspectral Images
by: Hou, Jinhui, et al.
Published: (2023)
by: Hou, Jinhui, et al.
Published: (2023)
An Efficient Quality Metric for Video Frame Interpolation Based on Motion-Field Divergence
by: Daly, Conall, et al.
Published: (2025)
by: Daly, Conall, et al.
Published: (2025)
SAM2S: Segment Anything in Surgical Videos via Semantic Long-term Tracking
by: Liu, Haofeng, et al.
Published: (2025)
by: Liu, Haofeng, et al.
Published: (2025)
A Novel Hybrid Approach for Retinal Vessel Segmentation with Dynamic Long-Range Dependency and Multi-Scale Retinal Edge Fusion Enhancement
by: Ouyang, Yihao, et al.
Published: (2025)
by: Ouyang, Yihao, et al.
Published: (2025)
Data-Efficient Learning for Generalizable Surgical Video Understanding
by: Nasirihaghighi, Sahar
Published: (2025)
by: Nasirihaghighi, Sahar
Published: (2025)
A Single-Frame and Multi-Frame Cascaded Image Super-Resolution Method
by: Sun, Jing, et al.
Published: (2024)
by: Sun, Jing, et al.
Published: (2024)
Inter-Frame Compression for Dynamic Point Cloud Geometry Coding
by: Akhtar, Anique, et al.
Published: (2022)
by: Akhtar, Anique, et al.
Published: (2022)
Goal-Oriented Semantic Communication for Wireless Visual Question Answering
by: Liu, Sige, et al.
Published: (2024)
by: Liu, Sige, et al.
Published: (2024)
DDUNet: Dual Dynamic U-Net for Highly-Efficient Cloud Segmentation
by: Li, Yijie, et al.
Published: (2025)
by: Li, Yijie, et al.
Published: (2025)
Similar Items
-
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
by: Wang, Shaoguang, et al.
Published: (2026) -
Event Camera Demosaicing via Swin Transformer and Pixel-focus Loss
by: Lu, Yunfan, et al.
Published: (2024) -
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
by: He, Jianxiang, et al.
Published: (2025) -
DrawVideo: Generating Long Video from Storyboard Keyframe Sketches
by: Xu, Chuanzhi, et al.
Published: (2026) -
Less is More: Token-Efficient Video-QA via Adaptive Frame-Pruning and Semantic Graph Integration
by: Wang, Shaoguang, et al.
Published: (2025)