Salvato in:
| Autori principali: | Edwards, Nicholas, Lee, Yukyung, Mao, Yujun Audrey, Qin, Yulu, Schuster, Sebastian, Kim, Najoung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2506.22598 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Code Pretraining Improves Entity Tracking Abilities of Language Models
di: Kim, Najoung, et al.
Pubblicazione: (2024)
di: Kim, Najoung, et al.
Pubblicazione: (2024)
CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklists
di: Lee, Yukyung, et al.
Pubblicazione: (2024)
di: Lee, Yukyung, et al.
Pubblicazione: (2024)
Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents
di: Edwards, Nicholas, et al.
Pubblicazione: (2026)
di: Edwards, Nicholas, et al.
Pubblicazione: (2026)
Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter It
di: Qin, Yulu, et al.
Pubblicazione: (2025)
di: Qin, Yulu, et al.
Pubblicazione: (2025)
A systematic framework for generating novel experimental hypotheses from language models
di: Misra, Kanishka, et al.
Pubblicazione: (2024)
di: Misra, Kanishka, et al.
Pubblicazione: (2024)
Do Language Models Track Entities Across State Changes?
di: Tang, Zilu, et al.
Pubblicazione: (2026)
di: Tang, Zilu, et al.
Pubblicazione: (2026)
Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document Streams
di: Lee, Yukyung, et al.
Pubblicazione: (2026)
di: Lee, Yukyung, et al.
Pubblicazione: (2026)
Is analogy enough to draw novel adjective-noun inferences?
di: Ross, Hayley, et al.
Pubblicazione: (2025)
di: Ross, Hayley, et al.
Pubblicazione: (2025)
Syn-QA2: Evaluating False Assumptions in Long-tail Questions with Synthetic QA Datasets
di: Daswani, Ashwin, et al.
Pubblicazione: (2024)
di: Daswani, Ashwin, et al.
Pubblicazione: (2024)
Is artificial intelligence still intelligence? LLMs generalize to novel adjective-noun pairs, but don't mimic the full human distribution
di: Ross, Hayley, et al.
Pubblicazione: (2024)
di: Ross, Hayley, et al.
Pubblicazione: (2024)
Can LLM-Augmented autonomous agents cooperate?, An evaluation of their cooperative capabilities through Melting Pot
di: Mosquera, Manuel, et al.
Pubblicazione: (2024)
di: Mosquera, Manuel, et al.
Pubblicazione: (2024)
RedacBench: Can AI Erase Your Secrets?
di: Jeon, Hyunjun, et al.
Pubblicazione: (2026)
di: Jeon, Hyunjun, et al.
Pubblicazione: (2026)
Death of the Novel(ty): Beyond n-Gram Novelty as a Metric for Textual Creativity
di: Saakyan, Arkadiy, et al.
Pubblicazione: (2025)
di: Saakyan, Arkadiy, et al.
Pubblicazione: (2025)
CIRF: Tokenizing Chain-of-Thoughts into Reusable Functional Units for Efficient Latent Reasoning in Large Language Models
di: Lee, Yukyung, et al.
Pubblicazione: (2026)
di: Lee, Yukyung, et al.
Pubblicazione: (2026)
Personas as a Way to Model Truthfulness in Language Models
di: Joshi, Nitish, et al.
Pubblicazione: (2023)
di: Joshi, Nitish, et al.
Pubblicazione: (2023)
CHAMP: A Competition-level Dataset for Fine-Grained Analyses of LLMs' Mathematical Reasoning Capabilities
di: Mao, Yujun, et al.
Pubblicazione: (2024)
di: Mao, Yujun, et al.
Pubblicazione: (2024)
Navigating the Path of Writing: Outline-guided Text Generation with Large Language Models
di: Lee, Yukyung, et al.
Pubblicazione: (2024)
di: Lee, Yukyung, et al.
Pubblicazione: (2024)
RareBench: Can LLMs Serve as Rare Diseases Specialists?
di: Chen, Xuanzhong, et al.
Pubblicazione: (2024)
di: Chen, Xuanzhong, et al.
Pubblicazione: (2024)
The why, what, and how of AI-based coding in scientific research
di: Zhuang, Tonghe, et al.
Pubblicazione: (2024)
di: Zhuang, Tonghe, et al.
Pubblicazione: (2024)
AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?
di: Ok, Hyunjong, et al.
Pubblicazione: (2025)
di: Ok, Hyunjong, et al.
Pubblicazione: (2025)
ClawBench: Can AI Agents Complete Everyday Online Tasks?
di: Zhang, Yuxuan, et al.
Pubblicazione: (2026)
di: Zhang, Yuxuan, et al.
Pubblicazione: (2026)
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
di: Kim, Eunsu, et al.
Pubblicazione: (2025)
di: Kim, Eunsu, et al.
Pubblicazione: (2025)
A systematic investigation of learnability from single child linguistic input
di: Qin, Yulu, et al.
Pubblicazione: (2024)
di: Qin, Yulu, et al.
Pubblicazione: (2024)
ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers?
di: Ye, Christine, et al.
Pubblicazione: (2025)
di: Ye, Christine, et al.
Pubblicazione: (2025)
REPRO-Bench: Can Agentic AI Systems Assess the Reproducibility of Social Science Research?
di: Hu, Chuxuan, et al.
Pubblicazione: (2025)
di: Hu, Chuxuan, et al.
Pubblicazione: (2025)
SysBench: Can Large Language Models Follow System Messages?
di: Qin, Yanzhao, et al.
Pubblicazione: (2024)
di: Qin, Yanzhao, et al.
Pubblicazione: (2024)
Diffusion LLMs can think EoS-by-EoS
di: Breckner, Sarah, et al.
Pubblicazione: (2026)
di: Breckner, Sarah, et al.
Pubblicazione: (2026)
Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
di: Prandi, Matteo, et al.
Pubblicazione: (2025)
di: Prandi, Matteo, et al.
Pubblicazione: (2025)
LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench
di: Valmeekam, Karthik, et al.
Pubblicazione: (2024)
di: Valmeekam, Karthik, et al.
Pubblicazione: (2024)
Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks
di: Wu, Zhaofeng, et al.
Pubblicazione: (2023)
di: Wu, Zhaofeng, et al.
Pubblicazione: (2023)
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
di: Son, Guijin, et al.
Pubblicazione: (2023)
di: Son, Guijin, et al.
Pubblicazione: (2023)
A Gradient Accumulation Method for Dense Retriever under Memory Constraint
di: Kim, Jaehee, et al.
Pubblicazione: (2024)
di: Kim, Jaehee, et al.
Pubblicazione: (2024)
Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models
di: Arad, Dana, et al.
Pubblicazione: (2025)
di: Arad, Dana, et al.
Pubblicazione: (2025)
A Behavior Tree-inspired programming language for autonomous agents
di: Biggar, Oliver, et al.
Pubblicazione: (2024)
di: Biggar, Oliver, et al.
Pubblicazione: (2024)
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
di: Kim, Haechan, et al.
Pubblicazione: (2026)
di: Kim, Haechan, et al.
Pubblicazione: (2026)
Can large language models build causal graphs?
di: Long, Stephanie, et al.
Pubblicazione: (2023)
di: Long, Stephanie, et al.
Pubblicazione: (2023)
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
di: Deng, Xiang, et al.
Pubblicazione: (2025)
di: Deng, Xiang, et al.
Pubblicazione: (2025)
CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models
di: Lee, Sangin, et al.
Pubblicazione: (2026)
di: Lee, Sangin, et al.
Pubblicazione: (2026)
FilBench: Can LLMs Understand and Generate Filipino?
di: Miranda, Lester James V., et al.
Pubblicazione: (2025)
di: Miranda, Lester James V., et al.
Pubblicazione: (2025)
On the Consideration of AI Openness: Can Good Intent Be Abused?
di: Kim, Yeeun, et al.
Pubblicazione: (2024)
di: Kim, Yeeun, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Code Pretraining Improves Entity Tracking Abilities of Language Models
di: Kim, Najoung, et al.
Pubblicazione: (2024) -
CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklists
di: Lee, Yukyung, et al.
Pubblicazione: (2024) -
Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents
di: Edwards, Nicholas, et al.
Pubblicazione: (2026) -
Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter It
di: Qin, Yulu, et al.
Pubblicazione: (2025) -
A systematic framework for generating novel experimental hypotheses from language models
di: Misra, Kanishka, et al.
Pubblicazione: (2024)