JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zhecan, Liu, Junzhang, Tang, Chia-Wei, Alomari, Hani, Sivakumar, Anushka, Sun, Rui, Li, Wenhao, Atabuzzaman, Md., Ayyubi, Hammad, You, Haoxuan, Ishmam, Alvi, Chang, Kai-Wei, Chang, Shih-Fu, Thomas, Chris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
von: Liu, Junzhang, et al.
Veröffentlicht: (2024)
von: Liu, Junzhang, et al.
Veröffentlicht: (2024)
ENTER: Event Based Interpretable Reasoning for VideoQA
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location Prediction
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
von: You, Haoxuan, et al.
Veröffentlicht: (2023)
von: You, Haoxuan, et al.
Veröffentlicht: (2023)
Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
von: Alomari, Hani, et al.
Veröffentlicht: (2025)
von: Alomari, Hani, et al.
Veröffentlicht: (2025)
UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding
von: Wang, Zhecan, et al.
Veröffentlicht: (2023)
von: Wang, Zhecan, et al.
Veröffentlicht: (2023)
RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos
von: Zare, Ali, et al.
Veröffentlicht: (2024)
von: Zare, Ali, et al.
Veröffentlicht: (2024)
Semantic Shield: Defending Vision-Language Models Against Backdooring and Poisoning via Fine-grained Knowledge Alignment
von: Ishmam, Alvi Md, et al.
Veröffentlicht: (2024)
von: Ishmam, Alvi Md, et al.
Veröffentlicht: (2024)
Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models
von: Atabuzzaman, Md., et al.
Veröffentlicht: (2025)
von: Atabuzzaman, Md., et al.
Veröffentlicht: (2025)
Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models
von: Atabuzzaman, Md., et al.
Veröffentlicht: (2025)
von: Atabuzzaman, Md., et al.
Veröffentlicht: (2025)
LAMP: Learning Universal Adversarial Perturbations for Multi-Image Tasks via Pre-trained Models
von: Ishmam, Alvi Md, et al.
Veröffentlicht: (2026)
von: Ishmam, Alvi Md, et al.
Veröffentlicht: (2026)
SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models
von: Hussain, Aafiya, et al.
Veröffentlicht: (2026)
von: Hussain, Aafiya, et al.
Veröffentlicht: (2026)
SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
von: Sivakumar, Anushka, et al.
Veröffentlicht: (2025)
von: Sivakumar, Anushka, et al.
Veröffentlicht: (2025)
Flexible-length Text Infilling for Discrete Diffusion Models
von: Zhang, Andrew, et al.
Veröffentlicht: (2025)
von: Zhang, Andrew, et al.
Veröffentlicht: (2025)
Video Summarization: Towards Entity-Aware Captions
von: Ayyubi, Hammad A., et al.
Veröffentlicht: (2023)
von: Ayyubi, Hammad A., et al.
Veröffentlicht: (2023)
Arabic Sentiment Analysis with Noisy Deep Explainable Model
von: Atabuzzaman, Md., et al.
Veröffentlicht: (2023)
von: Atabuzzaman, Md., et al.
Veröffentlicht: (2023)
What Matters in Explanations: Towards Explainable Fake Review Detection Focusing on Transformers
von: Shajalal, Md, et al.
Veröffentlicht: (2024)
von: Shajalal, Md, et al.
Veröffentlicht: (2024)
TimeWarp: Evaluating Web Agents by Revisiting the Past
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2026)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2026)
Generalized Choi-Davis-Jensen's Operator Inequalities and Their Applications
von: Chang, Shih Yu, et al.
Veröffentlicht: (2024)
von: Chang, Shih Yu, et al.
Veröffentlicht: (2024)
Textual Entailment Recognition with Semantic Features from Empirical Text Representation
von: Shajalal, Md, et al.
Veröffentlicht: (2022)
von: Shajalal, Md, et al.
Veröffentlicht: (2022)
Sporadic Hepatocellular Carcinoma in an Adolescent Harboring a Somatic CDH1 Mutation: A Case Report and Literature Review
von: Chang‐Wei Su, et al.
Veröffentlicht: (2026)
von: Chang‐Wei Su, et al.
Veröffentlicht: (2026)
FourierKAN outperforms MLP on Text Classification Head Fine-tuning
von: Imran, Abdullah Al, et al.
Veröffentlicht: (2024)
von: Imran, Abdullah Al, et al.
Veröffentlicht: (2024)
Recursive Optimal Stopping with Poisson Stopping Constraints
von: Liang, Gechun, et al.
Veröffentlicht: (2024)
von: Liang, Gechun, et al.
Veröffentlicht: (2024)
Don't Stop Believing: Mapping Distance Learners' Research Journeys
von: Brahme, Maria E., et al.
Veröffentlicht: (2016)
von: Brahme, Maria E., et al.
Veröffentlicht: (2016)
Zero-Shot Vehicle Model Recognition via Text-Based Retrieval-Augmented Generation
von: Chang, Wei-Chia, et al.
Veröffentlicht: (2025)
von: Chang, Wei-Chia, et al.
Veröffentlicht: (2025)
Visual Robustness Benchmark for Visual Question Answering (VQA)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
Efficient Ta 3 N 5 Photoanodes via Interface Engineering of Bixbyite‐Type Ta 2 N 3 Precursors
von: Chia‐Wei Chang, et al.
Veröffentlicht: (2025)
von: Chia‐Wei Chang, et al.
Veröffentlicht: (2025)
DORA Explorer: Improving the Exploration Ability of LLMs Without Training
von: Gurjar, Priya, et al.
Veröffentlicht: (2026)
von: Gurjar, Priya, et al.
Veröffentlicht: (2026)
Lift and leading-edge suction parameter of separated flows over an NACA0012 at high angles of attack
von: Chang, Ching, et al.
Veröffentlicht: (2026)
von: Chang, Ching, et al.
Veröffentlicht: (2026)
OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation
von: Chang, Jingjing, et al.
Veröffentlicht: (2025)
von: Chang, Jingjing, et al.
Veröffentlicht: (2025)
RinQ: Towards predicting central sites in proteins on current quantum computers
von: Mohtashim, Shah Ishmam
Veröffentlicht: (2025)
von: Mohtashim, Shah Ishmam
Veröffentlicht: (2025)
Merit Network Telescope: Processing and Initial Insights from Nearly 20 Years of Darknet Traffic for Cybersecurity Research
von: Ismail, Shereen, et al.
Veröffentlicht: (2025)
von: Ismail, Shereen, et al.
Veröffentlicht: (2025)
Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning
von: Sun, Renliang, et al.
Veröffentlicht: (2025)
von: Sun, Renliang, et al.
Veröffentlicht: (2025)
Tensor-Lifted Multivariate Functional Calculus Beyond Commutativity and Boundedness
von: Chang, Shih-Yu
Veröffentlicht: (2026)
von: Chang, Shih-Yu
Veröffentlicht: (2026)
Generalized Multiple Operator Integrals for Operators with Finite Dimensions
von: Chang, Shih-Yu
Veröffentlicht: (2025)
von: Chang, Shih-Yu
Veröffentlicht: (2025)
Existence and Enumeration of Polynomially Transformed Matrices under Spectral and Nilpotent Constraints
von: Chang, Shih-Yu
Veröffentlicht: (2025)
von: Chang, Shih-Yu
Veröffentlicht: (2025)
Spectral and Nilpotent Matrix Orderings: Comparison and Applications in Dynamic Systems
von: Chang, Shih-Yu
Veröffentlicht: (2025)
von: Chang, Shih-Yu
Veröffentlicht: (2025)
The Operadic Spectrum and Obstructions to Spectral Base Change
von: Chang, Shih-Yu
Veröffentlicht: (2026)
von: Chang, Shih-Yu
Veröffentlicht: (2026)
Multivariate Mond-Pecaric Method with Applications to Hypercomplex Function Sobolev Embedding
von: Chang, Shih-Yu
Veröffentlicht: (2024)
von: Chang, Shih-Yu
Veröffentlicht: (2024)
Categorified Spectral Sheaves and Homotopical Invariants for Noncommuting Operators
von: Chang, Shih-Yu
Veröffentlicht: (2026)
von: Chang, Shih-Yu
Veröffentlicht: (2026)
Ähnliche Einträge
-
Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
von: Liu, Junzhang, et al.
Veröffentlicht: (2024) -
ENTER: Event Based Interpretable Reasoning for VideoQA
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025) -
PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location Prediction
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025) -
IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
von: You, Haoxuan, et al.
Veröffentlicht: (2023) -
Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
von: Alomari, Hani, et al.
Veröffentlicht: (2025)