TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Zeqing, Zhang, Shiyuan, Tang, Chengpei, Wang, Keze |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EarthVL: A Progressive Earth Vision-Language Understanding and Generation Framework
por: Wang, Junjue, et al.
Publicado: (2026)
por: Wang, Junjue, et al.
Publicado: (2026)
DVGBench: Implicit-to-Explicit Visual Grounding Benchmark in UAV Imagery with Large Vision-Language Models
por: Zhou, Yue, et al.
Publicado: (2026)
por: Zhou, Yue, et al.
Publicado: (2026)
Skeletonization-Based Adversarial Perturbations on Large Vision Language Model's Mathematical Text Recognition
por: Yoshida, Masatomo, et al.
Publicado: (2026)
por: Yoshida, Masatomo, et al.
Publicado: (2026)
DVLA-RL: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning
por: Li, Wenhao, et al.
Publicado: (2026)
por: Li, Wenhao, et al.
Publicado: (2026)
DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response
por: Wang, Junjue, et al.
Publicado: (2025)
por: Wang, Junjue, et al.
Publicado: (2025)
Evaluating the Significance of Outdoor Advertising from Driver's Perspective Using Computer Vision
por: Černeková, Zuzana, et al.
Publicado: (2023)
por: Černeková, Zuzana, et al.
Publicado: (2023)
Exploring Diffusion with Test-Time Training on Efficient Image Restoration
por: Lu, Rongchang, et al.
Publicado: (2025)
por: Lu, Rongchang, et al.
Publicado: (2025)
Video-Based Human Pose Regression via Decoupled Space-Time Aggregation
por: He, Jijie, et al.
Publicado: (2024)
por: He, Jijie, et al.
Publicado: (2024)
TDIP: Tunable Deep Image Processing, a Real Time Melt Pool Monitoring Solution
por: Akhavan, Javid, et al.
Publicado: (2024)
por: Akhavan, Javid, et al.
Publicado: (2024)
MCA-Bench: A Multimodal Benchmark for Evaluating CAPTCHA Robustness Against VLM-based Attacks
por: Wu, Zonglin, et al.
Publicado: (2025)
por: Wu, Zonglin, et al.
Publicado: (2025)
SAM Encoder Breach by Adversarial Simplicial Complex Triggers Downstream Model Failures
por: Qin, Yi, et al.
Publicado: (2025)
por: Qin, Yi, et al.
Publicado: (2025)
SkeletonX: Data-Efficient Skeleton-based Action Recognition via Cross-sample Feature Aggregation
por: Zhang, Zongye, et al.
Publicado: (2025)
por: Zhang, Zongye, et al.
Publicado: (2025)
CigTime: Corrective Instruction Generation Through Inverse Motion Editing
por: Fang, Qihang, et al.
Publicado: (2024)
por: Fang, Qihang, et al.
Publicado: (2024)
Efficient Diffusion Models: A Comprehensive Survey from Principles to Practices
por: Ma, Zhiyuan, et al.
Publicado: (2024)
por: Ma, Zhiyuan, et al.
Publicado: (2024)
MdaIF: Robust One-Stop Multi-Degradation-Aware Image Fusion with Language-Driven Semantics
por: Li, Jing, et al.
Publicado: (2025)
por: Li, Jing, et al.
Publicado: (2025)
KNN Transformer with Pyramid Prompts for Few-Shot Learning
por: Li, Wenhao, et al.
Publicado: (2024)
por: Li, Wenhao, et al.
Publicado: (2024)
HyperFM: An Efficient Hyperspectral Foundation Model with Spectral Grouping
por: Tushar, Zahid Hassan, et al.
Publicado: (2026)
por: Tushar, Zahid Hassan, et al.
Publicado: (2026)
AVadCLIP: Audio-Visual Collaboration for Robust Video Anomaly Detection
por: Wu, Peng, et al.
Publicado: (2025)
por: Wu, Peng, et al.
Publicado: (2025)
Cross-View-Prediction: Exploring Contrastive Feature for Hyperspectral Image Classification
por: Zhang, Anyu, et al.
Publicado: (2022)
por: Zhang, Anyu, et al.
Publicado: (2022)
CMAB: A First National-Scale Multi-Attribute Building Dataset in China Derived from Open Source Data and GeoAI
por: Zhang, Yecheng, et al.
Publicado: (2024)
por: Zhang, Yecheng, et al.
Publicado: (2024)
VT-FSL: Bridging Vision and Text with LLMs for Few-Shot Learning
por: Li, Wenhao, et al.
Publicado: (2025)
por: Li, Wenhao, et al.
Publicado: (2025)
Robust Multi-Source Covid-19 Detection in CT Images
por: Pritha, Asmita Yuki, et al.
Publicado: (2026)
por: Pritha, Asmita Yuki, et al.
Publicado: (2026)
Rapid Adaptation of Earth Observation Foundation Models for Segmentation
por: Selvam, Karthick Panner, et al.
Publicado: (2024)
por: Selvam, Karthick Panner, et al.
Publicado: (2024)
Scalable and Realistic Virtual Try-on Application for Foundation Makeup with Kubelka-Munk Theory
por: Pang, Hui, et al.
Publicado: (2025)
por: Pang, Hui, et al.
Publicado: (2025)
Cost Savings from Automatic Quality Assessment of Generated Images
por: Giro-i-Nieto, Xavier, et al.
Publicado: (2025)
por: Giro-i-Nieto, Xavier, et al.
Publicado: (2025)
Tiny-YOLOSAM: Fast Hybrid Image Segmentation
por: Xu, Kenneth, et al.
Publicado: (2025)
por: Xu, Kenneth, et al.
Publicado: (2025)
A Deep Learning Approach to Identify Rock Bolts in Complex 3D Point Clouds of Underground Mines Captured Using Mobile Laser Scanners
por: Patra, Dibyayan, et al.
Publicado: (2025)
por: Patra, Dibyayan, et al.
Publicado: (2025)
SETR: A Two-Stage Semantic-Enhanced Framework for Zero-Shot Composed Image Retrieval
por: Xiao, Yuqi, et al.
Publicado: (2025)
por: Xiao, Yuqi, et al.
Publicado: (2025)
Supersampling of Data from Structured-light Scanner with Deep Learning
por: Melicherčík, Martin, et al.
Publicado: (2023)
por: Melicherčík, Martin, et al.
Publicado: (2023)
Group Activity Recognition using Unreliable Tracked Pose
por: Thilakarathne, Haritha, et al.
Publicado: (2024)
por: Thilakarathne, Haritha, et al.
Publicado: (2024)
Deepfake Detection Generalization with Diffusion Noise
por: Qi, Hongyuan, et al.
Publicado: (2026)
por: Qi, Hongyuan, et al.
Publicado: (2026)
Towards Integrated Rock Support Visualisation in 3D Point Cloud of Underground Mines
por: Patra, Dibyayan, et al.
Publicado: (2026)
por: Patra, Dibyayan, et al.
Publicado: (2026)
Processing and Segmentation of Human Teeth from 2D Images using Weakly Supervised Learning
por: Kunzo, Tomáš, et al.
Publicado: (2023)
por: Kunzo, Tomáš, et al.
Publicado: (2023)
EUFCC-340K: A Faceted Hierarchical Dataset for Metadata Annotation in GLAM Collections
por: Net, Francesc, et al.
Publicado: (2024)
por: Net, Francesc, et al.
Publicado: (2024)
Facial Spatiotemporal Graphs: Leveraging the 3D Facial Surface for Remote Physiological Measurement
por: Cantrill, Sam, et al.
Publicado: (2026)
por: Cantrill, Sam, et al.
Publicado: (2026)
Automated Discontinuity Set Characterisation in Enclosed Rock Face Point Clouds Using Single-Shot Filtering and Cyclic Orientation Transformation
por: Patra, Dibyayan, et al.
Publicado: (2026)
por: Patra, Dibyayan, et al.
Publicado: (2026)
Orientation-conditioned Facial Texture Mapping for Video-based Facial Remote Photoplethysmography Estimation
por: Cantrill, Sam, et al.
Publicado: (2024)
por: Cantrill, Sam, et al.
Publicado: (2024)
Dynamic Brightness Adaptation for Robust Multi-modal Image Fusion
por: Sun, Yiming, et al.
Publicado: (2024)
por: Sun, Yiming, et al.
Publicado: (2024)
On-the-Fly Guidance Training for Medical Image Registration
por: Xin, Yuelin, et al.
Publicado: (2023)
por: Xin, Yuelin, et al.
Publicado: (2023)
VersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis
por: Chen, Zhipeng, et al.
Publicado: (2024)
por: Chen, Zhipeng, et al.
Publicado: (2024)
Ejemplares similares
-
EarthVL: A Progressive Earth Vision-Language Understanding and Generation Framework
por: Wang, Junjue, et al.
Publicado: (2026) -
DVGBench: Implicit-to-Explicit Visual Grounding Benchmark in UAV Imagery with Large Vision-Language Models
por: Zhou, Yue, et al.
Publicado: (2026) -
Skeletonization-Based Adversarial Perturbations on Large Vision Language Model's Mathematical Text Recognition
por: Yoshida, Masatomo, et al.
Publicado: (2026) -
DVLA-RL: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning
por: Li, Wenhao, et al.
Publicado: (2026) -
DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response
por: Wang, Junjue, et al.
Publicado: (2025)