ASCIIEval: Benchmarking Models' Visual Perception in Text Strings via ASCII Art
Fuente:
arXiv
Salvato in:
| Autori principali: | Jia, Qi, Yue, Xiang, Huang, Shanshan, Qin, Ziheng, Liu, Yizhu, Lin, Bill Yuchen, You, Yang, Zhai, Guangtao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SimulBench: Evaluating Language Models with Creative Simulation Tasks
di: Jia, Qi, et al.
Pubblicazione: (2024)
di: Jia, Qi, et al.
Pubblicazione: (2024)
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
di: Jiang, Fengqing, et al.
Pubblicazione: (2024)
di: Jiang, Fengqing, et al.
Pubblicazione: (2024)
ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
di: Yang, Guan-Yan, et al.
Pubblicazione: (2025)
di: Yang, Guan-Yan, et al.
Pubblicazione: (2025)
Testing the Depth of ChatGPT's Comprehension via Cross-Modal Tasks Based on ASCII-Art: GPT3.5's Abilities in Regard to Recognizing and Generating ASCII-Art Are Not Totally Lacking
di: Bayani, David
Pubblicazione: (2023)
di: Bayani, David
Pubblicazione: (2023)
Information Density Principle for MLLM Benchmarks
di: Li, Chunyi, et al.
Pubblicazione: (2025)
di: Li, Chunyi, et al.
Pubblicazione: (2025)
Boosting LLM via Learning from Data Iteratively and Selectively
di: Jia, Qi, et al.
Pubblicazione: (2024)
di: Jia, Qi, et al.
Pubblicazione: (2024)
VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
di: Liu, Junpeng, et al.
Pubblicazione: (2024)
di: Liu, Junpeng, et al.
Pubblicazione: (2024)
Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
di: Song, Yifan, et al.
Pubblicazione: (2024)
di: Song, Yifan, et al.
Pubblicazione: (2024)
Meta-Chunking: Learning Text Segmentation and Semantic Completion via Logical Perception
di: Zhao, Jihao, et al.
Pubblicazione: (2024)
di: Zhao, Jihao, et al.
Pubblicazione: (2024)
Movie101v2: Improved Movie Narration Benchmark
di: Yue, Zihao, et al.
Pubblicazione: (2024)
di: Yue, Zihao, et al.
Pubblicazione: (2024)
One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework
di: Jia, Qi, et al.
Pubblicazione: (2025)
di: Jia, Qi, et al.
Pubblicazione: (2025)
MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis
di: Zhou, Yingjie, et al.
Pubblicazione: (2024)
di: Zhou, Yingjie, et al.
Pubblicazione: (2024)
TIGERScore: Towards Building Explainable Metric for All Text Generation Tasks
di: Jiang, Dongfu, et al.
Pubblicazione: (2023)
di: Jiang, Dongfu, et al.
Pubblicazione: (2023)
EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
di: Shen, Ye, et al.
Pubblicazione: (2026)
di: Shen, Ye, et al.
Pubblicazione: (2026)
TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency
di: Wang, Juntong, et al.
Pubblicazione: (2025)
di: Wang, Juntong, et al.
Pubblicazione: (2025)
Evading Toxicity Detection with ASCII-art: A Benchmark of Spatial Attacks on Moderation Systems
di: Berezin, Sergey, et al.
Pubblicazione: (2024)
di: Berezin, Sergey, et al.
Pubblicazione: (2024)
LitVISTA: A Benchmark for Narrative Orchestration in Literary Text
di: Lu, Mingzhe, et al.
Pubblicazione: (2026)
di: Lu, Mingzhe, et al.
Pubblicazione: (2026)
Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QA
di: Zhang, Kaiwei, et al.
Pubblicazione: (2025)
di: Zhang, Kaiwei, et al.
Pubblicazione: (2025)
LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition
di: Huang, Chengsong, et al.
Pubblicazione: (2023)
di: Huang, Chengsong, et al.
Pubblicazione: (2023)
SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
di: Zhu, Xiangyang, et al.
Pubblicazione: (2025)
di: Zhu, Xiangyang, et al.
Pubblicazione: (2025)
Statistical Analysis of Sentence Structures through ASCII, Lexical Alignment and PCA
di: Sahdev, Abhijeet
Pubblicazione: (2025)
di: Sahdev, Abhijeet
Pubblicazione: (2025)
OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation
di: Wang, Zilong, et al.
Pubblicazione: (2024)
di: Wang, Zilong, et al.
Pubblicazione: (2024)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
di: Xu, Zhangchen, et al.
Pubblicazione: (2024)
di: Xu, Zhangchen, et al.
Pubblicazione: (2024)
QoNext: Towards Next-generation QoE for Foundation Models
di: Guo, Yijin, et al.
Pubblicazione: (2025)
di: Guo, Yijin, et al.
Pubblicazione: (2025)
User-centric Subjective Leaderboard by Customizable Reward Modeling
di: Jia, Qi, et al.
Pubblicazione: (2025)
di: Jia, Qi, et al.
Pubblicazione: (2025)
Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs
di: Wang, Junying, et al.
Pubblicazione: (2025)
di: Wang, Junying, et al.
Pubblicazione: (2025)
Evaluating from Benign to Dynamic Adversarial: A Squid Game for Large Language Models
di: Chen, Zijian, et al.
Pubblicazione: (2025)
di: Chen, Zijian, et al.
Pubblicazione: (2025)
CULTURE-GEN: Revealing Global Cultural Perception in Language Models through Natural Language Prompting
di: Li, Huihan, et al.
Pubblicazione: (2024)
di: Li, Huihan, et al.
Pubblicazione: (2024)
A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
di: Shen, Ye, et al.
Pubblicazione: (2025)
di: Shen, Ye, et al.
Pubblicazione: (2025)
Stateful Evidence-Driven Retrieval-Augmented Generation with Iterative Reasoning
di: Dong, Qi, et al.
Pubblicazione: (2026)
di: Dong, Qi, et al.
Pubblicazione: (2026)
Are AI-Generated Text Detectors Robust to Adversarial Perturbations?
di: Huang, Guanhua, et al.
Pubblicazione: (2024)
di: Huang, Guanhua, et al.
Pubblicazione: (2024)
The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
di: Song, Yifan, et al.
Pubblicazione: (2024)
di: Song, Yifan, et al.
Pubblicazione: (2024)
Affordance Benchmark for MLLMs
di: Wang, Junying, et al.
Pubblicazione: (2025)
di: Wang, Junying, et al.
Pubblicazione: (2025)
MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models
di: Huang, Zhongzhan, et al.
Pubblicazione: (2025)
di: Huang, Zhongzhan, et al.
Pubblicazione: (2025)
Knowledge Fusion via Bidirectional Information Aggregation
di: Zhai, Songlin, et al.
Pubblicazione: (2025)
di: Zhai, Songlin, et al.
Pubblicazione: (2025)
Redundancy Principles for MLLMs Benchmarks
di: Zhang, Zicheng, et al.
Pubblicazione: (2025)
di: Zhang, Zicheng, et al.
Pubblicazione: (2025)
OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
di: Zheng, Tianyu, et al.
Pubblicazione: (2024)
di: Zheng, Tianyu, et al.
Pubblicazione: (2024)
PEPPER: Perception-Guided Perturbation for Robust Backdoor Defense in Text-to-Image Diffusion Models
di: Chew, Oscar, et al.
Pubblicazione: (2025)
di: Chew, Oscar, et al.
Pubblicazione: (2025)
Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models
di: Wang, Zhaochen, et al.
Pubblicazione: (2025)
di: Wang, Zhaochen, et al.
Pubblicazione: (2025)
Teaching LMMs for Image Quality Scoring and Interpreting
di: Zhang, Zicheng, et al.
Pubblicazione: (2025)
di: Zhang, Zicheng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SimulBench: Evaluating Language Models with Creative Simulation Tasks
di: Jia, Qi, et al.
Pubblicazione: (2024) -
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
di: Jiang, Fengqing, et al.
Pubblicazione: (2024) -
ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
di: Yang, Guan-Yan, et al.
Pubblicazione: (2025) -
Testing the Depth of ChatGPT's Comprehension via Cross-Modal Tasks Based on ASCII-Art: GPT3.5's Abilities in Regard to Recognizing and Generating ASCII-Art Are Not Totally Lacking
di: Bayani, David
Pubblicazione: (2023) -
Information Density Principle for MLLM Benchmarks
di: Li, Chunyi, et al.
Pubblicazione: (2025)