Defeasible Visual Entailment: Benchmark, Evaluator, and Reward-Driven Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yue, Jing, Liqiang, Gogate, Vibhav |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Video Large Multimodal Models Think Like Doubters-or Double-Down: A Study on Defeasible Video Entailment
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
Deep Dependency Networks and Advanced Inference Schemes for Multi-Label Classification
von: Arya, Shivvrat, et al.
Veröffentlicht: (2024)
von: Arya, Shivvrat, et al.
Veröffentlicht: (2024)
VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment
von: Saravanan, Darshana, et al.
Veröffentlicht: (2024)
von: Saravanan, Darshana, et al.
Veröffentlicht: (2024)
Compositional Entailment Learning for Hyperbolic Vision-Language Models
von: Pal, Avik, et al.
Veröffentlicht: (2024)
von: Pal, Avik, et al.
Veröffentlicht: (2024)
Towards Scene Graph Anticipation
von: Peddi, Rohith, et al.
Veröffentlicht: (2024)
von: Peddi, Rohith, et al.
Veröffentlicht: (2024)
Segment Any Vehicle: Semantic and Visual Context Driven SAM and A Benchmark
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Advancing Egocentric Video Question Answering with Multimodal Large Language Models
von: Patel, Alkesh, et al.
Veröffentlicht: (2025)
von: Patel, Alkesh, et al.
Veröffentlicht: (2025)
FuRL: Visual-Language Models as Fuzzy Rewards for Reinforcement Learning
von: Fu, Yuwei, et al.
Veröffentlicht: (2024)
von: Fu, Yuwei, et al.
Veröffentlicht: (2024)
A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
Gradient-Free Noise Optimization for Reward Alignment in Generative Models
von: Kim, Jeongsol, et al.
Veröffentlicht: (2026)
von: Kim, Jeongsol, et al.
Veröffentlicht: (2026)
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
von: Xiao, Tong, et al.
Veröffentlicht: (2025)
von: Xiao, Tong, et al.
Veröffentlicht: (2025)
Subtask-Aware Visual Reward Learning from Segmented Demonstrations
von: Kim, Changyeon, et al.
Veröffentlicht: (2025)
von: Kim, Changyeon, et al.
Veröffentlicht: (2025)
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
von: Chintapatla, Ishant, et al.
Veröffentlicht: (2025)
von: Chintapatla, Ishant, et al.
Veröffentlicht: (2025)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
von: Gao, Ziqi, et al.
Veröffentlicht: (2024)
von: Gao, Ziqi, et al.
Veröffentlicht: (2024)
A Large-scale Medical Visual Task Adaptation Benchmark
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
von: Ji, Haonian, et al.
Veröffentlicht: (2025)
von: Ji, Haonian, et al.
Veröffentlicht: (2025)
TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation
von: Gong, Han, et al.
Veröffentlicht: (2026)
von: Gong, Han, et al.
Veröffentlicht: (2026)
V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
von: Mai, Zheda, et al.
Veröffentlicht: (2025)
von: Mai, Zheda, et al.
Veröffentlicht: (2025)
Improving Prototypical Visual Explanations with Reward Reweighing, Reselection, and Retraining
von: Li, Aaron J., et al.
Veröffentlicht: (2023)
von: Li, Aaron J., et al.
Veröffentlicht: (2023)
Reward Score Matching: Unifying Reward-based Fine-tuning for Flow and Diffusion Models
von: Lee, Jeongjae, et al.
Veröffentlicht: (2026)
von: Lee, Jeongjae, et al.
Veröffentlicht: (2026)
DAVE: Diagnostic benchmark for Audio Visual Evaluation
von: Radevski, Gorjan, et al.
Veröffentlicht: (2025)
von: Radevski, Gorjan, et al.
Veröffentlicht: (2025)
MaRVL-QA: A Benchmark for Mathematical Reasoning over Visual Landscapes
von: Pande, Nilay, et al.
Veröffentlicht: (2025)
von: Pande, Nilay, et al.
Veröffentlicht: (2025)
MMS-VPR: Multimodal Street-Level Visual Place Recognition Dataset and Benchmark
von: Ou, Yiwei, et al.
Veröffentlicht: (2025)
von: Ou, Yiwei, et al.
Veröffentlicht: (2025)
Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models
von: Moayeri, Mazda, et al.
Veröffentlicht: (2024)
von: Moayeri, Mazda, et al.
Veröffentlicht: (2024)
Event-Driven Neuromorphic Vision Enables Energy-Efficient Visual Place Recognition
von: Keime, Geoffroy, et al.
Veröffentlicht: (2026)
von: Keime, Geoffroy, et al.
Veröffentlicht: (2026)
MVR: Multi-view Video Reward Shaping for Reinforcement Learning
von: Luo, Lirui, et al.
Veröffentlicht: (2026)
von: Luo, Lirui, et al.
Veröffentlicht: (2026)
Pretrained Reversible Generation as Unsupervised Visual Representation Learning
von: Xue, Rongkun, et al.
Veröffentlicht: (2024)
von: Xue, Rongkun, et al.
Veröffentlicht: (2024)
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
DKDM: Data-Free Knowledge Distillation for Diffusion Models with Any Architecture
von: Xiang, Qianlong, et al.
Veröffentlicht: (2024)
von: Xiang, Qianlong, et al.
Veröffentlicht: (2024)
Rethinking the Evaluation Protocol of Domain Generalization
von: Yu, Han, et al.
Veröffentlicht: (2023)
von: Yu, Han, et al.
Veröffentlicht: (2023)
RL for Consistency Models: Faster Reward Guided Text-to-Image Generation
von: Oertell, Owen, et al.
Veröffentlicht: (2024)
von: Oertell, Owen, et al.
Veröffentlicht: (2024)
Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts
von: Pan, Hongkun, et al.
Veröffentlicht: (2026)
von: Pan, Hongkun, et al.
Veröffentlicht: (2026)
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits
von: Yosef, Ron, et al.
Veröffentlicht: (2025)
von: Yosef, Ron, et al.
Veröffentlicht: (2025)
A-I-RAVEN and I-RAVEN-Mesh: Two New Benchmarks for Abstract Visual Reasoning
von: Małkiński, Mikołaj, et al.
Veröffentlicht: (2024)
von: Małkiński, Mikołaj, et al.
Veröffentlicht: (2024)
Synthetic History: Evaluating Visual Representations of the Past in Diffusion Models
von: Palmini, Maria-Teresa De Rosa, et al.
Veröffentlicht: (2025)
von: Palmini, Maria-Teresa De Rosa, et al.
Veröffentlicht: (2025)
Diff-Instruct with Diffused Reward: Towards Principled One-step Generator RL
von: Wu, Junyi, et al.
Veröffentlicht: (2026)
von: Wu, Junyi, et al.
Veröffentlicht: (2026)
COVR:Collaborative Optimization of VLMs and RL Agent for Visual-Based Control
von: Xia, Canming, et al.
Veröffentlicht: (2026)
von: Xia, Canming, et al.
Veröffentlicht: (2026)
Hierarchical Variational Policies for Reward-Guided Diffusion
von: Pandey, Kushagra, et al.
Veröffentlicht: (2026)
von: Pandey, Kushagra, et al.
Veröffentlicht: (2026)
Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Can Video Large Multimodal Models Think Like Doubters-or Double-Down: A Study on Defeasible Video Entailment
von: Zhang, Yue, et al.
Veröffentlicht: (2025) -
Deep Dependency Networks and Advanced Inference Schemes for Multi-Label Classification
von: Arya, Shivvrat, et al.
Veröffentlicht: (2024) -
VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment
von: Saravanan, Darshana, et al.
Veröffentlicht: (2024) -
Compositional Entailment Learning for Hyperbolic Vision-Language Models
von: Pal, Avik, et al.
Veröffentlicht: (2024) -
Towards Scene Graph Anticipation
von: Peddi, Rohith, et al.
Veröffentlicht: (2024)