AAAR-1.0: Assessing AI's Potential to Assist Research
Fuente:
arXiv
Saved in:
| Main Authors: | Lou, Renze, Xu, Hanzi, Wang, Sijia, Du, Jiangshu, Kamoi, Ryo, Lu, Xiaoxin, Xie, Jian, Sun, Yuxuan, Zhang, Yusen, Ahn, Jihyun Janice, Fang, Hongchao, Zou, Zhuoyang, Ma, Wenchao, Li, Xi, Zhang, Kai, Xia, Congying, Huang, Lifu, Yin, Wenpeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Direct-Inverse Prompting: Analyzing LLMs' Discriminative Capacity in Self-Improving Generation
by: Ahn, Jihyun Janice, et al.
Published: (2024)
by: Ahn, Jihyun Janice, et al.
Published: (2024)
MUFFIN: Curating Multi-Faceted Instructions for Improving Instruction-Following
by: Lou, Renze, et al.
Published: (2023)
by: Lou, Renze, et al.
Published: (2023)
Evaluating LLMs at Detecting Errors in LLM Responses
by: Kamoi, Ryo, et al.
Published: (2024)
by: Kamoi, Ryo, et al.
Published: (2024)
Prompt-Reverse Inconsistency: LLM Self-Inconsistency Beyond Generative Randomness and Prompt Paraphrasing
by: Ahn, Jihyun Janice, et al.
Published: (2025)
by: Ahn, Jihyun Janice, et al.
Published: (2025)
LLMs' Classification Performance is Overclaimed
by: Xu, Hanzi, et al.
Published: (2024)
by: Xu, Hanzi, et al.
Published: (2024)
Large Language Models for Mathematical Reasoning: Progresses and Challenges
by: Ahn, Janice, et al.
Published: (2024)
by: Ahn, Janice, et al.
Published: (2024)
Toward Zero-Shot Instruction Following
by: Lou, Renze, et al.
Published: (2023)
by: Lou, Renze, et al.
Published: (2023)
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?
by: Zhang, Yusen, et al.
Published: (2025)
by: Zhang, Yusen, et al.
Published: (2025)
Large Language Model Instruction Following: A Survey of Progresses and Challenges
by: Lou, Renze, et al.
Published: (2023)
by: Lou, Renze, et al.
Published: (2023)
When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs
by: Kamoi, Ryo, et al.
Published: (2024)
by: Kamoi, Ryo, et al.
Published: (2024)
Efficient PRM Training Data Synthesis via Formal Verification
by: Kamoi, Ryo, et al.
Published: (2025)
by: Kamoi, Ryo, et al.
Published: (2025)
X-Shot: A Unified System to Handle Frequent, Few-shot and Zero-shot Learning Simultaneously in Classification
by: Xu, Hanzi, et al.
Published: (2024)
by: Xu, Hanzi, et al.
Published: (2024)
ScaleFormer: Span Representation Cumulation for Long-Context Transformer
by: Du, Jiangshu, et al.
Published: (2025)
by: Du, Jiangshu, et al.
Published: (2025)
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
by: Li, Xi, et al.
Published: (2024)
by: Li, Xi, et al.
Published: (2024)
Could AI Trace and Explain the Origins of AI-Generated Images and Text?
by: Fang, Hongchao, et al.
Published: (2025)
by: Fang, Hongchao, et al.
Published: (2025)
VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
by: Kamoi, Ryo, et al.
Published: (2024)
by: Kamoi, Ryo, et al.
Published: (2024)
Can Prompt Modifiers Control Bias? A Comparative Analysis of Text-to-Image Generative Models
by: Shin, Philip Wootaek, et al.
Published: (2024)
by: Shin, Philip Wootaek, et al.
Published: (2024)
FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability
by: Xia, Congying, et al.
Published: (2024)
by: Xia, Congying, et al.
Published: (2024)
GReaTer: Gradients over Reasoning Makes Smaller Language Models Strong Prompt Optimizers
by: Das, Sarkar Snigdha Sarathi, et al.
Published: (2024)
by: Das, Sarkar Snigdha Sarathi, et al.
Published: (2024)
Fair Abstractive Summarization of Diverse Perspectives
by: Zhang, Yusen, et al.
Published: (2023)
by: Zhang, Yusen, et al.
Published: (2023)
M2-Verify: A Large-Scale Multidomain Benchmark for Checking Multimodal Claim Consistency
by: Ansari, Abolfazl, et al.
Published: (2026)
by: Ansari, Abolfazl, et al.
Published: (2026)
DIAGPaper: Diagnosing Valid and Specific Weaknesses in Scientific Papers via Multi-Agent Reasoning
by: Zou, Zhuoyang, et al.
Published: (2026)
by: Zou, Zhuoyang, et al.
Published: (2026)
Debate as Optimization: Adaptive Conformal Prediction and Diverse Retrieval for Event Extraction
by: Wang, Sijia, et al.
Published: (2024)
by: Wang, Sijia, et al.
Published: (2024)
Targeted Augmentation for Low-Resource Event Extraction
by: Wang, Sijia, et al.
Published: (2024)
by: Wang, Sijia, et al.
Published: (2024)
Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation
by: Lu, Xiaoxin, et al.
Published: (2025)
by: Lu, Xiaoxin, et al.
Published: (2025)
Bridging the Know-Act Gap via Task-Level Autoregressive Reasoning
by: Ahn, Jihyun Janice, et al.
Published: (2026)
by: Ahn, Jihyun Janice, et al.
Published: (2026)
UMIE: Unified Multimodal Information Extraction with Instruction Tuning
by: Sun, Lin, et al.
Published: (2024)
by: Sun, Lin, et al.
Published: (2024)
The Tool Illusion: Rethinking Tool Use in Web Agents
by: Lou, Renze, et al.
Published: (2026)
by: Lou, Renze, et al.
Published: (2026)
Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction
by: Kamoi, Ryo, et al.
Published: (2026)
by: Kamoi, Ryo, et al.
Published: (2026)
MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge Conflicts
by: Xie, Jian, et al.
Published: (2023)
by: Xie, Jian, et al.
Published: (2023)
Die Ordnung des Theaters. Eine Soziologie der Regie
by: Hänzi, Denis
Published: (2015)
by: Hänzi, Denis
Published: (2015)
LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing
by: Du, Jiangshu, et al.
Published: (2024)
by: Du, Jiangshu, et al.
Published: (2024)
Some interesting number theory problems
by: Zhang, Wenpeng
Published: (2025)
by: Zhang, Wenpeng
Published: (2025)
A century problem related to the Legendre symbol modulo p
by: Zhang, Wenpeng
Published: (2025)
by: Zhang, Wenpeng
Published: (2025)
NeuroGen: Neural Network Parameter Generation via Large Language Models
by: Wang, Jiaqi, et al.
Published: (2025)
by: Wang, Jiaqi, et al.
Published: (2025)
Holographic Intelligence Surface Assisted Integrated Sensing and Communication
by: Liu, Zhuoyang, et al.
Published: (2024)
by: Liu, Zhuoyang, et al.
Published: (2024)
Advancing Chart Question Answering with Robust Chart Component Recognition
by: Zheng, Hanwen, et al.
Published: (2024)
by: Zheng, Hanwen, et al.
Published: (2024)
PyFi: Toward Pyramid-like Financial Image Understanding for VLMs via Adversarial Agents
by: Zhang, Yuqun, et al.
Published: (2025)
by: Zhang, Yuqun, et al.
Published: (2025)
A General Benchmark Framework is Dynamic Graph Neural Network Need
by: Zhang, Yusen
Published: (2024)
by: Zhang, Yusen
Published: (2024)
Similar Items
-
Direct-Inverse Prompting: Analyzing LLMs' Discriminative Capacity in Self-Improving Generation
by: Ahn, Jihyun Janice, et al.
Published: (2024) -
MUFFIN: Curating Multi-Faceted Instructions for Improving Instruction-Following
by: Lou, Renze, et al.
Published: (2023) -
Evaluating LLMs at Detecting Errors in LLM Responses
by: Kamoi, Ryo, et al.
Published: (2024) -
Prompt-Reverse Inconsistency: LLM Self-Inconsistency Beyond Generative Randomness and Prompt Paraphrasing
by: Ahn, Jihyun Janice, et al.
Published: (2025) -
LLMs' Classification Performance is Overclaimed
by: Xu, Hanzi, et al.
Published: (2024)