Evaluation Hallucination in Multi-Round Incomplete Information Lateral-Driven Reasoning Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Wenhan, Hu, Tianyi, Zheng, Jingyi, Sun, Zhen, Zhao, Yuemeng, Liu, Yule, He, Xinlei, Huang, Xinyi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
by: Zheng, Jingyi, et al.
Published: (2025)
by: Zheng, Jingyi, et al.
Published: (2025)
CHASM: Unveiling Covert Advertisements on Chinese Social Media
by: Zheng, Jingyi, et al.
Published: (2026)
by: Zheng, Jingyi, et al.
Published: (2026)
ZPD-SCA: Unveiling the Blind Spots of LLMs in Assessing Students' Cognitive Abilities
by: Dong, Wenhan, et al.
Published: (2025)
by: Dong, Wenhan, et al.
Published: (2025)
Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications
by: Dong, Wenhan, et al.
Published: (2025)
by: Dong, Wenhan, et al.
Published: (2025)
Quantized Delta Weight Is Safety Keeper
by: Liu, Yule, et al.
Published: (2024)
by: Liu, Yule, et al.
Published: (2024)
Thought Manipulation: External Thought Can Be Efficient for Large Reasoning Models
by: Liu, Yule, et al.
Published: (2025)
by: Liu, Yule, et al.
Published: (2025)
GasAgent: A Multi-Agent Framework for Automated Gas Optimization in Smart Contracts
by: Zheng, Jingyi, et al.
Published: (2025)
by: Zheng, Jingyi, et al.
Published: (2025)
JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
by: Peng, Zifan, et al.
Published: (2025)
by: Peng, Zifan, et al.
Published: (2025)
Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models
by: Zhang, Zongmin, et al.
Published: (2025)
by: Zhang, Zongmin, et al.
Published: (2025)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
by: Zheng, Jingyi, et al.
Published: (2024)
by: Zheng, Jingyi, et al.
Published: (2024)
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
by: Luo, Zeren, et al.
Published: (2025)
by: Luo, Zeren, et al.
Published: (2025)
Privacy-Preserving Federated Learning via Homomorphic Adversarial Networks
by: Dong, Wenhan, et al.
Published: (2024)
by: Dong, Wenhan, et al.
Published: (2024)
PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-Tuning
by: Sun, Zhen, et al.
Published: (2024)
by: Sun, Zhen, et al.
Published: (2024)
LatEval: An Interactive LLMs Evaluation Benchmark with Incomplete Information from Lateral Thinking Puzzles
by: Huang, Shulin, et al.
Published: (2023)
by: Huang, Shulin, et al.
Published: (2023)
"I Can See Forever!": Evaluating Real-time VideoLLMs for Assisting Individuals with Visual Impairments
by: Zhang, Ziyi, et al.
Published: (2025)
by: Zhang, Ziyi, et al.
Published: (2025)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
by: Zhang, Heyi, et al.
Published: (2025)
by: Zhang, Heyi, et al.
Published: (2025)
MedAtlas: Evaluating LLMs for Multi-Round, Multi-Task Medical Reasoning Across Diverse Imaging Modalities and Clinical Text
by: Xu, Ronghao, et al.
Published: (2025)
by: Xu, Ronghao, et al.
Published: (2025)
KnowGuard: Knowledge-Driven Abstention for Multi-Round Clinical Reasoning
by: Dang, Xilin, et al.
Published: (2025)
by: Dang, Xilin, et al.
Published: (2025)
RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
by: Sun, Wenfang, et al.
Published: (2026)
by: Sun, Wenfang, et al.
Published: (2026)
Stego Battlefield: Evaluating Image Steganography Attacks and Steganalysis Defenses
by: Sun, Zhen, et al.
Published: (2026)
by: Sun, Zhen, et al.
Published: (2026)
The Role of Information Incompleteness in Defending Against Stealth Attacks
by: Sun, Ke, et al.
Published: (2025)
by: Sun, Ke, et al.
Published: (2025)
On the Generalization and Adaptation Ability of Machine-Generated Text Detectors in Academic Writing
by: Liu, Yule, et al.
Published: (2024)
by: Liu, Yule, et al.
Published: (2024)
Comment on “Translating Chuang Tzu into world literature: text and context”
by: Yuemeng Ge
Published: (2023)
by: Yuemeng Ge
Published: (2023)
Comment on “How should we think about common prosperity and challenges in the context of financialization?”
by: Yuemeng Ge
Published: (2023)
by: Yuemeng Ge
Published: (2023)
Auditing Data Membership in Reinforcement Learning With Verifiable Rewards
by: Liu, Yule, et al.
Published: (2025)
by: Liu, Yule, et al.
Published: (2025)
Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media
by: Sun, Zhen, et al.
Published: (2024)
by: Sun, Zhen, et al.
Published: (2024)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
by: Yi, Sibo, et al.
Published: (2024)
by: Yi, Sibo, et al.
Published: (2024)
Chain-of-Authorization: Embedding authorization into large language models
by: Li, Yang, et al.
Published: (2026)
by: Li, Yang, et al.
Published: (2026)
The Multi-Round Diagnostic RAG Framework for Emulating Clinical Reasoning
by: Sun, Penglei, et al.
Published: (2025)
by: Sun, Penglei, et al.
Published: (2025)
VCR: Learning Valid Contextual Representation for Incomplete Wearable Signals
by: Weng, Yuxuan, et al.
Published: (2026)
by: Weng, Yuxuan, et al.
Published: (2026)
DAOpt: Modeling and Evaluation of Data-Driven Optimization under Uncertainty with LLMs
by: Zhu, WenZhuo, et al.
Published: (2025)
by: Zhu, WenZhuo, et al.
Published: (2025)
MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images
by: Tong, Qinyue, et al.
Published: (2025)
by: Tong, Qinyue, et al.
Published: (2025)
Visible‐Light‐Driven Sustainable Chemical Vapor Deposition of Polymer Films and Patterns
by: Yuanhao Shen, et al.
Published: (2025)
by: Yuanhao Shen, et al.
Published: (2025)
Information Recovery-Driven Deep Incomplete Multiview Clustering Network
by: Liu, Chengliang, et al.
Published: (2023)
by: Liu, Chengliang, et al.
Published: (2023)
Work Zones challenge VLM Trajectory Planning: Toward Mitigation and Robust Autonomous Driving
by: Liao, Yifan, et al.
Published: (2025)
by: Liao, Yifan, et al.
Published: (2025)
Hyperbolic Enhanced Representation Learning for Incomplete Multi-view Clustering
by: Chen, Tianyi, et al.
Published: (2026)
by: Chen, Tianyi, et al.
Published: (2026)
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
by: Ran, Delong, et al.
Published: (2024)
by: Ran, Delong, et al.
Published: (2024)
Structure-Grounded Knowledge Retrieval via Code Dependencies for Multi-Step Data Reasoning
by: Huang, Xinyi
Published: (2026)
by: Huang, Xinyi
Published: (2026)
MedMAP: Promoting Incomplete Multi-modal Brain Tumor Segmentation with Alignment
by: Liu, Tianyi, et al.
Published: (2024)
by: Liu, Tianyi, et al.
Published: (2024)
Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks
by: Yang, Cheng, et al.
Published: (2025)
by: Yang, Cheng, et al.
Published: (2025)
Similar Items
-
TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
by: Zheng, Jingyi, et al.
Published: (2025) -
CHASM: Unveiling Covert Advertisements on Chinese Social Media
by: Zheng, Jingyi, et al.
Published: (2026) -
ZPD-SCA: Unveiling the Blind Spots of LLMs in Assessing Students' Cognitive Abilities
by: Dong, Wenhan, et al.
Published: (2025) -
Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications
by: Dong, Wenhan, et al.
Published: (2025) -
Quantized Delta Weight Is Safety Keeper
by: Liu, Yule, et al.
Published: (2024)