DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math?
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Young-Suk, Astudillo, Ramon Fernandez, Florian, Radu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Multi-Document Grounded Multi-Turn Synthetic Dialog Generation
di: Lee, Young-Suk, et al.
Pubblicazione: (2024)
di: Lee, Young-Suk, et al.
Pubblicazione: (2024)
Self-Refinement of Language Models from External Proxy Metrics Feedback
di: Ramji, Keshav, et al.
Pubblicazione: (2024)
di: Ramji, Keshav, et al.
Pubblicazione: (2024)
Optimal Policy Minimum Bayesian Risk
di: Astudillo, Ramón Fernandez, et al.
Pubblicazione: (2025)
di: Astudillo, Ramón Fernandez, et al.
Pubblicazione: (2025)
Do LLMs Benefit From Their Own Words?
di: Huang, Jenny Y., et al.
Pubblicazione: (2026)
di: Huang, Jenny Y., et al.
Pubblicazione: (2026)
Do Agents Know What They Can't Do? Evaluating Feasibility Awareness in Tool-Using Agents
di: Cheng, Liang, et al.
Pubblicazione: (2026)
di: Cheng, Liang, et al.
Pubblicazione: (2026)
REGENT: A Retrieval-Augmented Generalist Agent That Can Act In-Context in New Environments
di: Sridhar, Kaustubh, et al.
Pubblicazione: (2024)
di: Sridhar, Kaustubh, et al.
Pubblicazione: (2024)
Latent Principle Discovery for Language Model Self-Improvement
di: Ramji, Keshav, et al.
Pubblicazione: (2025)
di: Ramji, Keshav, et al.
Pubblicazione: (2025)
DERA: Dense Entity Retrieval for Entity Alignment in Knowledge Graphs
di: Wang, Zhichun, et al.
Pubblicazione: (2024)
di: Wang, Zhichun, et al.
Pubblicazione: (2024)
CHILL at SemEval-2025 Task 2: You Can't Just Throw Entities and Hope -- Make Your LLM to Get Them Right
di: Lee, Jaebok, et al.
Pubblicazione: (2025)
di: Lee, Jaebok, et al.
Pubblicazione: (2025)
Case-Based or Rule-Based: How Do Transformers Do the Math?
di: Hu, Yi, et al.
Pubblicazione: (2024)
di: Hu, Yi, et al.
Pubblicazione: (2024)
Your LLM Agent Can Leak Your Data: Data Exfiltration via Backdoored Tool Use
di: Zhang, Wuyang, et al.
Pubblicazione: (2026)
di: Zhang, Wuyang, et al.
Pubblicazione: (2026)
Your Classifier Can Do More: Towards Balancing the Gaps in Classification, Robustness, and Generation
di: Jiang, Kaichao, et al.
Pubblicazione: (2025)
di: Jiang, Kaichao, et al.
Pubblicazione: (2025)
Re$^2$Math: Benchmarking Theorem Retrieval in Research-Level Mathematics
di: Lyu, Zicheng, et al.
Pubblicazione: (2026)
di: Lyu, Zicheng, et al.
Pubblicazione: (2026)
AgenticSimLaw: A Juvenile Courtroom Multi-Agent Debate Simulation for Explainable High-Stakes Tabular Decision Making
di: Chun, Jon, et al.
Pubblicazione: (2026)
di: Chun, Jon, et al.
Pubblicazione: (2026)
Can LLMs Master Math? Investigating Large Language Models on Math Stack Exchange
di: Satpute, Ankit, et al.
Pubblicazione: (2024)
di: Satpute, Ankit, et al.
Pubblicazione: (2024)
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
di: Yu, Longhui, et al.
Pubblicazione: (2023)
di: Yu, Longhui, et al.
Pubblicazione: (2023)
Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
di: Hu, Yijie, et al.
Pubblicazione: (2025)
di: Hu, Yijie, et al.
Pubblicazione: (2025)
Do Math Reasoning LLMs Help Predict the Impact of Public Transit Events?
di: Fang, Bowen, et al.
Pubblicazione: (2025)
di: Fang, Bowen, et al.
Pubblicazione: (2025)
Multi-Agent Architecture in Distributed Environment Control Systems: vision, challenges, and opportunities
di: Astudillo, Natasha, et al.
Pubblicazione: (2025)
di: Astudillo, Natasha, et al.
Pubblicazione: (2025)
Your Agent Can Defend Itself against Backdoor Attacks
di: Changjiang, Li, et al.
Pubblicazione: (2025)
di: Changjiang, Li, et al.
Pubblicazione: (2025)
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
Do Phone-Use Agents Respect Your Privacy?
di: Tang, Zhengyang, et al.
Pubblicazione: (2026)
di: Tang, Zhengyang, et al.
Pubblicazione: (2026)
Agent Manufacturing: Foundation-Model Agents as First-Class Industrial Entities
di: Zhang, Yilei
Pubblicazione: (2026)
di: Zhang, Yilei
Pubblicazione: (2026)
Your Code Agent Can Grow Alongside You with Structured Memory
di: Deng, Yi-Xuan, et al.
Pubblicazione: (2026)
di: Deng, Yi-Xuan, et al.
Pubblicazione: (2026)
Are Your Agents Upward Deceivers?
di: Guo, Dadi, et al.
Pubblicazione: (2025)
di: Guo, Dadi, et al.
Pubblicazione: (2025)
RealMath: A Continuous Benchmark for Evaluating Language Models on Research-Level Mathematics
di: Zhang, Jie, et al.
Pubblicazione: (2025)
di: Zhang, Jie, et al.
Pubblicazione: (2025)
Do not Abstain! Identify and Solve the Uncertainty
di: Liu, Jingyu, et al.
Pubblicazione: (2025)
di: Liu, Jingyu, et al.
Pubblicazione: (2025)
Models Can and Should Embrace the Communicative Nature of Human-Generated Math
di: Boguraev, Sasha, et al.
Pubblicazione: (2024)
di: Boguraev, Sasha, et al.
Pubblicazione: (2024)
Can AI Tools Transform Low-Demand Math Tasks? An Evaluation of Task Modification Capabilities
di: Fox, Danielle S., et al.
Pubblicazione: (2026)
di: Fox, Danielle S., et al.
Pubblicazione: (2026)
SenseMath: Do LLMs Have Number Sense? Evaluating Shortcut Use, Judgment, and Generation
di: Zhuang, Haomin, et al.
Pubblicazione: (2026)
di: Zhuang, Haomin, et al.
Pubblicazione: (2026)
StepMathAgent: A Step-Wise Agent for Evaluating Mathematical Processes through Tree-of-Error
di: Yang, Shu-Xun, et al.
Pubblicazione: (2025)
di: Yang, Shu-Xun, et al.
Pubblicazione: (2025)
From Text to Visuals: Using LLMs to Generate Math Diagrams with Vector Graphics
di: Lee, Jaewook, et al.
Pubblicazione: (2025)
di: Lee, Jaewook, et al.
Pubblicazione: (2025)
Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts?
di: Karim, Aabid, et al.
Pubblicazione: (2025)
di: Karim, Aabid, et al.
Pubblicazione: (2025)
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
di: Zhang, Renrui, et al.
Pubblicazione: (2024)
di: Zhang, Renrui, et al.
Pubblicazione: (2024)
Do Similar Entities have Similar Embeddings?
di: Hubert, Nicolas, et al.
Pubblicazione: (2023)
di: Hubert, Nicolas, et al.
Pubblicazione: (2023)
Do Language Models Track Entities Across State Changes?
di: Tang, Zilu, et al.
Pubblicazione: (2026)
di: Tang, Zilu, et al.
Pubblicazione: (2026)
Knowledge Tagging System on Math Questions via LLMs with Flexible Demonstration Retriever
di: Li, Hang, et al.
Pubblicazione: (2024)
di: Li, Hang, et al.
Pubblicazione: (2024)
Can We Predict Your Next Move Without Breaking Your Privacy?
di: Soni, Arpita, et al.
Pubblicazione: (2025)
di: Soni, Arpita, et al.
Pubblicazione: (2025)
AtomicRAG: Atom-Entity Graphs for Retrieval-Augmented Generation
di: Hou, Yanning, et al.
Pubblicazione: (2026)
di: Hou, Yanning, et al.
Pubblicazione: (2026)
Motion-to-Response Content Generation via Multi-Agent AI System with Real-Time Safety Verification
di: Lee, HyeYoung
Pubblicazione: (2026)
di: Lee, HyeYoung
Pubblicazione: (2026)
Documenti analoghi
-
Multi-Document Grounded Multi-Turn Synthetic Dialog Generation
di: Lee, Young-Suk, et al.
Pubblicazione: (2024) -
Self-Refinement of Language Models from External Proxy Metrics Feedback
di: Ramji, Keshav, et al.
Pubblicazione: (2024) -
Optimal Policy Minimum Bayesian Risk
di: Astudillo, Ramón Fernandez, et al.
Pubblicazione: (2025) -
Do LLMs Benefit From Their Own Words?
di: Huang, Jenny Y., et al.
Pubblicazione: (2026) -
Do Agents Know What They Can't Do? Evaluating Feasibility Awareness in Tool-Using Agents
di: Cheng, Liang, et al.
Pubblicazione: (2026)