VinaBench: Benchmark for Faithful and Consistent Visual Narratives
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Silin, Mathew, Sheryl, Mi, Li, Mamooler, Sepideh, Zhao, Mengjie, Wakaki, Hiromi, Mitsufuji, Yuki, Montariol, Syrielle, Bosselut, Antoine |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PICLe: Pseudo-Annotations for In-Context Learning in Low-Resource Named Entity Detection
by: Mamooler, Sepideh, et al.
Published: (2024)
by: Mamooler, Sepideh, et al.
Published: (2024)
DiffuCOMET: Contextual Commonsense Knowledge Diffusion
by: Gao, Silin, et al.
Published: (2024)
by: Gao, Silin, et al.
Published: (2024)
ComperDial: Commonsense Persona-grounded Dialogue Dataset and Benchmark
by: Wakaki, Hiromi, et al.
Published: (2024)
by: Wakaki, Hiromi, et al.
Published: (2024)
Rethinking Skill Extraction in the Job Market Domain using Large Language Models
by: Nguyen, Khanh Cao, et al.
Published: (2024)
by: Nguyen, Khanh Cao, et al.
Published: (2024)
Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios
by: Corbière, Charles, et al.
Published: (2025)
by: Corbière, Charles, et al.
Published: (2025)
JOBSKAPE: A Framework for Generating Synthetic Job Postings to Enhance Skill Matching
by: Magron, Antoine, et al.
Published: (2024)
by: Magron, Antoine, et al.
Published: (2024)
ConVQG: Contrastive Visual Question Generation with Multimodal Guidance
by: Mi, Li, et al.
Published: (2024)
by: Mi, Li, et al.
Published: (2024)
CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
by: Bhagwatkar, Rishika, et al.
Published: (2025)
by: Bhagwatkar, Rishika, et al.
Published: (2025)
"Flex Tape Can't Fix That": Bias and Misinformation in Edited Language Models
by: Halevy, Karina, et al.
Published: (2024)
by: Halevy, Karina, et al.
Published: (2024)
DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning
by: Mao, Zhuoyuan, et al.
Published: (2025)
by: Mao, Zhuoyuan, et al.
Published: (2025)
Instruction-tuning Aligns LLMs to the Human Brain
by: Aw, Khai Loong, et al.
Published: (2023)
by: Aw, Khai Loong, et al.
Published: (2023)
ConGeo: Robust Cross-view Geo-localization across Ground View Variations
by: Mi, Li, et al.
Published: (2024)
by: Mi, Li, et al.
Published: (2024)
Learning to Route Languages for Multilingual Policy Optimization
by: Guo, Geyang, et al.
Published: (2026)
by: Guo, Geyang, et al.
Published: (2026)
AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
by: Gao, Silin, et al.
Published: (2025)
by: Gao, Silin, et al.
Published: (2025)
Cross-Modal Learning for Music-to-Music-Video Description Generation
by: Mao, Zhuoyuan, et al.
Published: (2025)
by: Mao, Zhuoyuan, et al.
Published: (2025)
Course Recommender Systems Need to Consider the Job Market
by: Frej, Jibril, et al.
Published: (2024)
by: Frej, Jibril, et al.
Published: (2024)
CARE: Multilingual Human Preference Learning for Cultural Awareness
by: Guo, Geyang, et al.
Published: (2025)
by: Guo, Geyang, et al.
Published: (2025)
OpenMU: Your Swiss Army Knife for Music Understanding
by: Zhao, Mengjie, et al.
Published: (2024)
by: Zhao, Mengjie, et al.
Published: (2024)
MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
by: Wu, Qiyu, et al.
Published: (2025)
by: Wu, Qiyu, et al.
Published: (2025)
Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning
by: Kim, Kyuhee, et al.
Published: (2026)
by: Kim, Kyuhee, et al.
Published: (2026)
TED: Turn Emphasis with Dialogue Feature Attention for Emotion Recognition in Conversation
by: Ono, Junya, et al.
Published: (2025)
by: Ono, Junya, et al.
Published: (2025)
Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning
by: Paul, Debjit, et al.
Published: (2024)
by: Paul, Debjit, et al.
Published: (2024)
Intrinsic User-Centric Interpretability through Global Mixture of Experts
by: Swamy, Vinitra, et al.
Published: (2024)
by: Swamy, Vinitra, et al.
Published: (2024)
Using Natural Language Inference to Improve Persona Extraction from Dialogue in a New Domain
by: DeLucia, Alexandra, et al.
Published: (2024)
by: DeLucia, Alexandra, et al.
Published: (2024)
Checkmate: interpretable and explainable RSVQA is the endgame
by: Tosato, Lucrezia, et al.
Published: (2025)
by: Tosato, Lucrezia, et al.
Published: (2025)
Multi-Task Learning for Features Extraction in Financial Annual Reports
by: Montariol, Syrielle, et al.
Published: (2024)
by: Montariol, Syrielle, et al.
Published: (2024)
Demystifying MaskGIT Sampler and Beyond: Adaptive Order Selection in Masked Diffusion
by: Hayakawa, Satoshi, et al.
Published: (2025)
by: Hayakawa, Satoshi, et al.
Published: (2025)
Distillation of Discrete Diffusion through Dimensional Correlations
by: Hayakawa, Satoshi, et al.
Published: (2024)
by: Hayakawa, Satoshi, et al.
Published: (2024)
Few-shot Dialogue Strategy Learning for Motivational Interviewing via Inductive Reasoning
by: Xie, Zhouhang, et al.
Published: (2024)
by: Xie, Zhouhang, et al.
Published: (2024)
When AI Takes Sides on Questions of Faith: Persistent Asymmetries in AI-Mediated Faith Guidance
by: Israelsen, Brett, et al.
Published: (2026)
by: Israelsen, Brett, et al.
Published: (2026)
Reliable Evaluation and Benchmarks for Statement Autoformalization
by: Poiroux, Auguste, et al.
Published: (2024)
by: Poiroux, Auguste, et al.
Published: (2024)
Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation
by: Yang, Shiqi, et al.
Published: (2024)
by: Yang, Shiqi, et al.
Published: (2024)
GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models
by: Mirza, M. Jehanzeb, et al.
Published: (2024)
by: Mirza, M. Jehanzeb, et al.
Published: (2024)
GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration
by: Mi, Li, et al.
Published: (2025)
by: Mi, Li, et al.
Published: (2025)
SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
by: Zhang, Chenkai, et al.
Published: (2025)
by: Zhang, Chenkai, et al.
Published: (2025)
Theoretical Refinement of CLIP by Utilizing Linear Structure of Optimal Similarity
by: Yoshida, Naoki, et al.
Published: (2025)
by: Yoshida, Naoki, et al.
Published: (2025)
RLMEval: Evaluating Research-Level Neural Theorem Proving
by: Poiroux, Auguste, et al.
Published: (2025)
by: Poiroux, Auguste, et al.
Published: (2025)
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs
by: Bao, Forrest Sheng, et al.
Published: (2024)
by: Bao, Forrest Sheng, et al.
Published: (2024)
VCT: Training Consistency Models with Variational Noise Coupling
by: Silvestri, Gianluigi, et al.
Published: (2025)
by: Silvestri, Gianluigi, et al.
Published: (2025)
LogicGaze: Benchmarking Causal Consistency in Visual Narratives via Counterfactual Verification
by: Driscoll, Rory, et al.
Published: (2026)
by: Driscoll, Rory, et al.
Published: (2026)
Similar Items
-
PICLe: Pseudo-Annotations for In-Context Learning in Low-Resource Named Entity Detection
by: Mamooler, Sepideh, et al.
Published: (2024) -
DiffuCOMET: Contextual Commonsense Knowledge Diffusion
by: Gao, Silin, et al.
Published: (2024) -
ComperDial: Commonsense Persona-grounded Dialogue Dataset and Benchmark
by: Wakaki, Hiromi, et al.
Published: (2024) -
Rethinking Skill Extraction in the Job Market Domain using Large Language Models
by: Nguyen, Khanh Cao, et al.
Published: (2024) -
Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios
by: Corbière, Charles, et al.
Published: (2025)