JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhecan, Liu, Junzhang, Tang, Chia-Wei, Alomari, Hani, Sivakumar, Anushka, Sun, Rui, Li, Wenhao, Atabuzzaman, Md., Ayyubi, Hammad, You, Haoxuan, Ishmam, Alvi, Chang, Kai-Wei, Chang, Shih-Fu, Thomas, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
by: Liu, Junzhang, et al.
Published: (2024)
by: Liu, Junzhang, et al.
Published: (2024)
ENTER: Event Based Interpretable Reasoning for VideoQA
by: Ayyubi, Hammad, et al.
Published: (2025)
by: Ayyubi, Hammad, et al.
Published: (2025)
PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location Prediction
by: Ayyubi, Hammad, et al.
Published: (2025)
by: Ayyubi, Hammad, et al.
Published: (2025)
IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
by: You, Haoxuan, et al.
Published: (2023)
by: You, Haoxuan, et al.
Published: (2023)
Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
by: Alomari, Hani, et al.
Published: (2025)
by: Alomari, Hani, et al.
Published: (2025)
UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding
by: Wang, Zhecan, et al.
Published: (2023)
by: Wang, Zhecan, et al.
Published: (2023)
RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos
by: Zare, Ali, et al.
Published: (2024)
by: Zare, Ali, et al.
Published: (2024)
Semantic Shield: Defending Vision-Language Models Against Backdooring and Poisoning via Fine-grained Knowledge Alignment
by: Ishmam, Alvi Md, et al.
Published: (2024)
by: Ishmam, Alvi Md, et al.
Published: (2024)
Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models
by: Atabuzzaman, Md., et al.
Published: (2025)
by: Atabuzzaman, Md., et al.
Published: (2025)
Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models
by: Atabuzzaman, Md., et al.
Published: (2025)
by: Atabuzzaman, Md., et al.
Published: (2025)
LAMP: Learning Universal Adversarial Perturbations for Multi-Image Tasks via Pre-trained Models
by: Ishmam, Alvi Md, et al.
Published: (2026)
by: Ishmam, Alvi Md, et al.
Published: (2026)
SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models
by: Hussain, Aafiya, et al.
Published: (2026)
by: Hussain, Aafiya, et al.
Published: (2026)
SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
by: Sivakumar, Anushka, et al.
Published: (2025)
by: Sivakumar, Anushka, et al.
Published: (2025)
Flexible-length Text Infilling for Discrete Diffusion Models
by: Zhang, Andrew, et al.
Published: (2025)
by: Zhang, Andrew, et al.
Published: (2025)
Video Summarization: Towards Entity-Aware Captions
by: Ayyubi, Hammad A., et al.
Published: (2023)
by: Ayyubi, Hammad A., et al.
Published: (2023)
Arabic Sentiment Analysis with Noisy Deep Explainable Model
by: Atabuzzaman, Md., et al.
Published: (2023)
by: Atabuzzaman, Md., et al.
Published: (2023)
What Matters in Explanations: Towards Explainable Fake Review Detection Focusing on Transformers
by: Shajalal, Md, et al.
Published: (2024)
by: Shajalal, Md, et al.
Published: (2024)
TimeWarp: Evaluating Web Agents by Revisiting the Past
by: Ishmam, Md Farhan, et al.
Published: (2026)
by: Ishmam, Md Farhan, et al.
Published: (2026)
Generalized Choi-Davis-Jensen's Operator Inequalities and Their Applications
by: Chang, Shih Yu, et al.
Published: (2024)
by: Chang, Shih Yu, et al.
Published: (2024)
Textual Entailment Recognition with Semantic Features from Empirical Text Representation
by: Shajalal, Md, et al.
Published: (2022)
by: Shajalal, Md, et al.
Published: (2022)
Sporadic Hepatocellular Carcinoma in an Adolescent Harboring a Somatic CDH1 Mutation: A Case Report and Literature Review
by: Chang‐Wei Su, et al.
Published: (2026)
by: Chang‐Wei Su, et al.
Published: (2026)
FourierKAN outperforms MLP on Text Classification Head Fine-tuning
by: Imran, Abdullah Al, et al.
Published: (2024)
by: Imran, Abdullah Al, et al.
Published: (2024)
Recursive Optimal Stopping with Poisson Stopping Constraints
by: Liang, Gechun, et al.
Published: (2024)
by: Liang, Gechun, et al.
Published: (2024)
Don't Stop Believing: Mapping Distance Learners' Research Journeys
by: Brahme, Maria E., et al.
Published: (2016)
by: Brahme, Maria E., et al.
Published: (2016)
Zero-Shot Vehicle Model Recognition via Text-Based Retrieval-Augmented Generation
by: Chang, Wei-Chia, et al.
Published: (2025)
by: Chang, Wei-Chia, et al.
Published: (2025)
Visual Robustness Benchmark for Visual Question Answering (VQA)
by: Ishmam, Md Farhan, et al.
Published: (2024)
by: Ishmam, Md Farhan, et al.
Published: (2024)
Efficient Ta 3 N 5 Photoanodes via Interface Engineering of Bixbyite‐Type Ta 2 N 3 Precursors
by: Chia‐Wei Chang, et al.
Published: (2025)
by: Chia‐Wei Chang, et al.
Published: (2025)
DORA Explorer: Improving the Exploration Ability of LLMs Without Training
by: Gurjar, Priya, et al.
Published: (2026)
by: Gurjar, Priya, et al.
Published: (2026)
Lift and leading-edge suction parameter of separated flows over an NACA0012 at high angles of attack
by: Chang, Ching, et al.
Published: (2026)
by: Chang, Ching, et al.
Published: (2026)
OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation
by: Chang, Jingjing, et al.
Published: (2025)
by: Chang, Jingjing, et al.
Published: (2025)
RinQ: Towards predicting central sites in proteins on current quantum computers
by: Mohtashim, Shah Ishmam
Published: (2025)
by: Mohtashim, Shah Ishmam
Published: (2025)
Merit Network Telescope: Processing and Initial Insights from Nearly 20 Years of Darknet Traffic for Cybersecurity Research
by: Ismail, Shereen, et al.
Published: (2025)
by: Ismail, Shereen, et al.
Published: (2025)
Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning
by: Sun, Renliang, et al.
Published: (2025)
by: Sun, Renliang, et al.
Published: (2025)
Tensor-Lifted Multivariate Functional Calculus Beyond Commutativity and Boundedness
by: Chang, Shih-Yu
Published: (2026)
by: Chang, Shih-Yu
Published: (2026)
Generalized Multiple Operator Integrals for Operators with Finite Dimensions
by: Chang, Shih-Yu
Published: (2025)
by: Chang, Shih-Yu
Published: (2025)
Existence and Enumeration of Polynomially Transformed Matrices under Spectral and Nilpotent Constraints
by: Chang, Shih-Yu
Published: (2025)
by: Chang, Shih-Yu
Published: (2025)
Spectral and Nilpotent Matrix Orderings: Comparison and Applications in Dynamic Systems
by: Chang, Shih-Yu
Published: (2025)
by: Chang, Shih-Yu
Published: (2025)
The Operadic Spectrum and Obstructions to Spectral Base Change
by: Chang, Shih-Yu
Published: (2026)
by: Chang, Shih-Yu
Published: (2026)
Multivariate Mond-Pecaric Method with Applications to Hypercomplex Function Sobolev Embedding
by: Chang, Shih-Yu
Published: (2024)
by: Chang, Shih-Yu
Published: (2024)
Categorified Spectral Sheaves and Homotopical Invariants for Noncommuting Operators
by: Chang, Shih-Yu
Published: (2026)
by: Chang, Shih-Yu
Published: (2026)
Similar Items
-
Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
by: Liu, Junzhang, et al.
Published: (2024) -
ENTER: Event Based Interpretable Reasoning for VideoQA
by: Ayyubi, Hammad, et al.
Published: (2025) -
PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location Prediction
by: Ayyubi, Hammad, et al.
Published: (2025) -
IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
by: You, Haoxuan, et al.
Published: (2023) -
Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
by: Alomari, Hani, et al.
Published: (2025)