Saved in:
| Main Authors: | Jang, Yunseok, Song, Yeda, Sohn, Sungryull, Logeswaran, Lajanugen, Luo, Tiange, Kim, Dong-Ki, Bae, Kyunghoon, Lee, Honglak |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.12632 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents
by: Fu, Yao, et al.
Published: (2024)
by: Fu, Yao, et al.
Published: (2024)
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
by: Kim, Jaekyeom, et al.
Published: (2024)
by: Kim, Jaekyeom, et al.
Published: (2024)
Visual Test-time Scaling for GUI Agent Grounding
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Scaling Web Agent Training through Automatic Data Generation and Fine-grained Evaluation
by: Logeswaran, Lajanugen, et al.
Published: (2026)
by: Logeswaran, Lajanugen, et al.
Published: (2026)
Selective LoRA for Visual Tokens and Attention Heads
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
by: Khalifa, Muhammad, et al.
Published: (2026)
by: Khalifa, Muhammad, et al.
Published: (2026)
Do Not Trust Licenses You See: Dataset Compliance Requires Massive-Scale AI-Powered Lifecycle Tracing
by: Kim, Jaekyeom, et al.
Published: (2025)
by: Kim, Jaekyeom, et al.
Published: (2025)
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
by: Khalifa, Muhammad, et al.
Published: (2023)
by: Khalifa, Muhammad, et al.
Published: (2023)
Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?
by: Shen, Siqi, et al.
Published: (2025)
by: Shen, Siqi, et al.
Published: (2025)
Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense
by: Shen, Siqi, et al.
Published: (2024)
by: Shen, Siqi, et al.
Published: (2024)
Small Language Models Need Strong Verifiers to Self-Correct Reasoning
by: Zhang, Yunxiang, et al.
Published: (2024)
by: Zhang, Yunxiang, et al.
Published: (2024)
Process Reward Models That Think
by: Khalifa, Muhammad, et al.
Published: (2025)
by: Khalifa, Muhammad, et al.
Published: (2025)
View Selection for 3D Captioning via Diffusion Ranking
by: Luo, Tiange, et al.
Published: (2024)
by: Luo, Tiange, et al.
Published: (2024)
Significantly improving zero-shot X-ray pathology classification via fine-tuning pre-trained image-text encoders
by: Jang, Jongseong, et al.
Published: (2022)
by: Jang, Jongseong, et al.
Published: (2022)
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
by: Zhang, Lechen, et al.
Published: (2025)
by: Zhang, Lechen, et al.
Published: (2025)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
by: Zheng, Mingqian, et al.
Published: (2023)
by: Zheng, Mingqian, et al.
Published: (2023)
SPRIG: Improving Large Language Model Performance by System Prompt Optimization
by: Zhang, Lechen, et al.
Published: (2024)
by: Zhang, Lechen, et al.
Published: (2024)
Interactive and Expressive Code-Augmented Planning with Large Language Models
by: Liu, Anthony Z., et al.
Published: (2024)
by: Liu, Anthony Z., et al.
Published: (2024)
Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks
by: Lee, Changho, et al.
Published: (2024)
by: Lee, Changho, et al.
Published: (2024)
LGAI-EMBEDDING-Preview Technical Report
by: Choi, Jooyoung, et al.
Published: (2025)
by: Choi, Jooyoung, et al.
Published: (2025)
Probing Visual Language Priors in VLMs
by: Luo, Tiange, et al.
Published: (2024)
by: Luo, Tiange, et al.
Published: (2024)
YTCommentQA: Video Question Answerability in Instructional Videos
by: Yang, Saelyne, et al.
Published: (2024)
by: Yang, Saelyne, et al.
Published: (2024)
MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
by: Zhang, Yunxiang, et al.
Published: (2025)
by: Zhang, Yunxiang, et al.
Published: (2025)
You don't need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments
by: Shu, Bangzhao, et al.
Published: (2023)
by: Shu, Bangzhao, et al.
Published: (2023)
Deep Exploration of Cross-Lingual Zero-Shot Generalization in Instruction Tuning
by: Han, Janghoon, et al.
Published: (2024)
by: Han, Janghoon, et al.
Published: (2024)
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
by: Kim, Geon-Hyeong, et al.
Published: (2025)
by: Kim, Geon-Hyeong, et al.
Published: (2025)
Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
by: Kim, Jaeho, et al.
Published: (2025)
by: Kim, Jaeho, et al.
Published: (2025)
FaceGCD: Generalized Face Discovery via Dynamic Prefix Generation
by: Oh, Yunseok, et al.
Published: (2025)
by: Oh, Yunseok, et al.
Published: (2025)
Multi-lingual Multi-institutional Electronic Health Record based Predictive Model
by: Hur, Kyunghoon, et al.
Published: (2026)
by: Hur, Kyunghoon, et al.
Published: (2026)
DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
by: Han, Janghoon, et al.
Published: (2025)
by: Han, Janghoon, et al.
Published: (2025)
RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning Models
by: Han, Yunseok, et al.
Published: (2026)
by: Han, Yunseok, et al.
Published: (2026)
Learning Video Temporal Dynamics with Cross-Modal Attention for Robust Audio-Visual Speech Recognition
by: Kim, Sungnyun, et al.
Published: (2024)
by: Kim, Sungnyun, et al.
Published: (2024)
ReConPatch : Contrastive Patch Representation Learning for Industrial Anomaly Detection
by: Hyun, Jeeho, et al.
Published: (2023)
by: Hyun, Jeeho, et al.
Published: (2023)
Cross-lingual QA: A Key to Unlocking In-context Cross-lingual Performance
by: Kim, Sunkyoung, et al.
Published: (2023)
by: Kim, Sunkyoung, et al.
Published: (2023)
Descriptive Image-Text Matching with Graded Contextual Similarity
by: Jang, Jinhyun, et al.
Published: (2025)
by: Jang, Jinhyun, et al.
Published: (2025)
Cross-Platform Video Person ReID: A New Benchmark Dataset and Adaptation Approach
by: Zhang, Shizhou, et al.
Published: (2024)
by: Zhang, Shizhou, et al.
Published: (2024)
Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
by: Hwang, Jinwoo, et al.
Published: (2025)
by: Hwang, Jinwoo, et al.
Published: (2025)
SEDD: Scalable and Efficient Dataset Deduplication with GPUs
by: Son, Youngjun, et al.
Published: (2025)
by: Son, Youngjun, et al.
Published: (2025)
Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling
by: Chu, Sanghyeok, et al.
Published: (2026)
by: Chu, Sanghyeok, et al.
Published: (2026)
Exploring Cross-Domain Few-Shot Classification via Frequency-Aware Prompting
by: Zhang, Tiange, et al.
Published: (2024)
by: Zhang, Tiange, et al.
Published: (2024)
Similar Items
-
AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents
by: Fu, Yao, et al.
Published: (2024) -
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
by: Kim, Jaekyeom, et al.
Published: (2024) -
Visual Test-time Scaling for GUI Agent Grounding
by: Luo, Tiange, et al.
Published: (2025) -
Scaling Web Agent Training through Automatic Data Generation and Fine-grained Evaluation
by: Logeswaran, Lajanugen, et al.
Published: (2026) -
Selective LoRA for Visual Tokens and Attention Heads
by: Luo, Tiange, et al.
Published: (2025)