BRIGHT+: Upgrading the BRIGHT Benchmark with MARCUS, a Multi-Agent RAG Clean-Up Suite
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Liyang, Cai, Yujun, Dong, Jieqiong, Wang, Yiwei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval
por: Su, Hongjin, et al.
Publicado: (2024)
por: Su, Hongjin, et al.
Publicado: (2024)
AudioRouter: Data Efficient Audio Understanding via RL based Dual Reasoning
por: Chen, Liyang, et al.
Publicado: (2026)
por: Chen, Liyang, et al.
Publicado: (2026)
BRIGHT: A globally distributed multimodal building damage assessment dataset with very-high-resolution for all-weather disaster response
por: Chen, Hongruixuan, et al.
Publicado: (2025)
por: Chen, Hongruixuan, et al.
Publicado: (2025)
BRIGHT STAR ASTROMETRY WITH URAT
por: N. Zacharias
Publicado: (2015)
por: N. Zacharias
Publicado: (2015)
Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression
por: Cui, Yu, et al.
Publicado: (2025)
por: Cui, Yu, et al.
Publicado: (2025)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
por: Fu, Honghao, et al.
Publicado: (2026)
por: Fu, Honghao, et al.
Publicado: (2026)
Cure or Poison? Embedding Instructions Visually Alters Hallucination in Vision-Language Models
por: Wang, Zhaochen, et al.
Publicado: (2025)
por: Wang, Zhaochen, et al.
Publicado: (2025)
Process or Result? Manipulated Ending Tokens Can Mislead Reasoning LLMs to Ignore the Correct Reasoning Steps
por: Cui, Yu, et al.
Publicado: (2025)
por: Cui, Yu, et al.
Publicado: (2025)
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
por: Wu, Yike, et al.
Publicado: (2025)
por: Wu, Yike, et al.
Publicado: (2025)
MM-BRIGHT: A Multi-Task Multimodal Benchmark for Reasoning-Intensive Retrieval
por: Abdallah, Abdelrahman, et al.
Publicado: (2026)
por: Abdallah, Abdelrahman, et al.
Publicado: (2026)
An Executable Benchmarking Suite for Tool-Using Agents
por: Zhong, Zhiqing, et al.
Publicado: (2026)
por: Zhong, Zhiqing, et al.
Publicado: (2026)
Self-Manager: Parallel Agent Loop for Long-form Deep Research
por: Xu, Yilong, et al.
Publicado: (2026)
por: Xu, Yilong, et al.
Publicado: (2026)
AudioMotionBench: Evaluating Auditory Motion Perception in Audio LLMs
por: Sun, Zhe, et al.
Publicado: (2025)
por: Sun, Zhe, et al.
Publicado: (2025)
MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
por: Ge, Haonan, et al.
Publicado: (2025)
por: Ge, Haonan, et al.
Publicado: (2025)
Primacy Effect of ChatGPT
por: Wang, Yiwei, et al.
Publicado: (2023)
por: Wang, Yiwei, et al.
Publicado: (2023)
MARCUS: An agentic, multimodal vision-language model for cardiac diagnosis and management
por: O'Sullivan, Jack W, et al.
Publicado: (2026)
por: O'Sullivan, Jack W, et al.
Publicado: (2026)
THE BRIGHT FUTURE OF THE ASTROPHYSICAL MAGNETIC FIELD RESEARCH
por: A. Lazarian
Publicado: (2009)
por: A. Lazarian
Publicado: (2009)
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
por: Bragg, Jonathan, et al.
Publicado: (2025)
por: Bragg, Jonathan, et al.
Publicado: (2025)
RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation
por: Wu, Hang, et al.
Publicado: (2025)
por: Wu, Hang, et al.
Publicado: (2025)
Revisiting RAG Ensemble: A Theoretical and Mechanistic Analysis of Multi-RAG System Collaboration
por: Chen, Yifei, et al.
Publicado: (2025)
por: Chen, Yifei, et al.
Publicado: (2025)
COMMA: A Communicative Multimodal Multi-Agent Benchmark
por: Ossowski, Timothy, et al.
Publicado: (2024)
por: Ossowski, Timothy, et al.
Publicado: (2024)
EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design
por: Molinari, Gioele, et al.
Publicado: (2026)
por: Molinari, Gioele, et al.
Publicado: (2026)
BRIGHT PROSPECTS OR AN OMINOUS FUTURE. Anticipating Oil in Uganda
por: Annika Witte
Publicado: (2017)
por: Annika Witte
Publicado: (2017)
MAESTRO: Multi-Agent Evaluation Suite for Testing, Reliability, and Observability
por: Ma, Tie, et al.
Publicado: (2026)
por: Ma, Tie, et al.
Publicado: (2026)
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
por: Lam, Man Ho, et al.
Publicado: (2026)
por: Lam, Man Ho, et al.
Publicado: (2026)
ContextNav: Towards Agentic Multimodal In-Context Learning
por: Fu, Honghao, et al.
Publicado: (2025)
por: Fu, Honghao, et al.
Publicado: (2025)
MultiTab: A Comprehensive Benchmark Suite for Multi-Dimensional Evaluation in Tabular Domains
por: Lee, Kyungeun, et al.
Publicado: (2025)
por: Lee, Kyungeun, et al.
Publicado: (2025)
BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics
por: Fa, Dionizije, et al.
Publicado: (2026)
por: Fa, Dionizije, et al.
Publicado: (2026)
MoDeSuite: Robot Learning Task Suite for Benchmarking Mobile Manipulation with Deformable Objects
por: Zhang, Yuying, et al.
Publicado: (2025)
por: Zhang, Yuying, et al.
Publicado: (2025)
MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents
por: Wang, Luyuan, et al.
Publicado: (2024)
por: Wang, Luyuan, et al.
Publicado: (2024)
WavefrontDiffusion: Dynamic Decoding Schedule for Improved Reasoning
por: Yang, Haojin, et al.
Publicado: (2025)
por: Yang, Haojin, et al.
Publicado: (2025)
Enhancing LLM Character-Level Manipulation via Divide and Conquer
por: Xiong, Zhen, et al.
Publicado: (2025)
por: Xiong, Zhen, et al.
Publicado: (2025)
A MYSTERIOUS UNIVERSE REVEALING THE BRIGHT AND DARK SIDES OF THE COSMOS
por: Susana Planelles
Publicado: (2017)
por: Susana Planelles
Publicado: (2017)
PA-RAG: RAG Alignment via Multi-Perspective Preference Optimization
por: Wu, Jiayi, et al.
Publicado: (2024)
por: Wu, Jiayi, et al.
Publicado: (2024)
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
por: Sun, Bowen, et al.
Publicado: (2025)
por: Sun, Bowen, et al.
Publicado: (2025)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
por: Ge, Haonan, et al.
Publicado: (2025)
por: Ge, Haonan, et al.
Publicado: (2025)
Experience as a Compass: Multi-agent RAG with Evolving Orchestration and Agent Prompts
por: Li, Sha, et al.
Publicado: (2026)
por: Li, Sha, et al.
Publicado: (2026)
BRIGHT: A Collaborative Generalist-Specialist Foundation Model for Breast Pathology
por: Guo, Xiaojing, et al.
Publicado: (2026)
por: Guo, Xiaojing, et al.
Publicado: (2026)
Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
por: Yeh, Samuel, et al.
Publicado: (2025)
por: Yeh, Samuel, et al.
Publicado: (2025)
VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft
por: Fu, Honghao, et al.
Publicado: (2025)
por: Fu, Honghao, et al.
Publicado: (2025)
Ejemplares similares
-
BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval
por: Su, Hongjin, et al.
Publicado: (2024) -
AudioRouter: Data Efficient Audio Understanding via RL based Dual Reasoning
por: Chen, Liyang, et al.
Publicado: (2026) -
BRIGHT: A globally distributed multimodal building damage assessment dataset with very-high-resolution for all-weather disaster response
por: Chen, Hongruixuan, et al.
Publicado: (2025) -
BRIGHT STAR ASTROMETRY WITH URAT
por: N. Zacharias
Publicado: (2015) -
Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression
por: Cui, Yu, et al.
Publicado: (2025)