Guardado en:
| Autores principales: | Song, Sangjun, Oh, Minjae, Lee, Seungkyu, Jo, Sungmin, Jo, Yohan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2510.00546 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
KL for a KL: On-Policy Distillation with Control Variate Baseline
por: Oh, Minjae, et al.
Publicado: (2026)
por: Oh, Minjae, et al.
Publicado: (2026)
In-N-Out: A Parameter-Level API Graph Dataset for Tool Agents
por: Lee, Seungkyu, et al.
Publicado: (2025)
por: Lee, Seungkyu, et al.
Publicado: (2025)
Future Policy Approximation for Offline Reinforcement Learning Improves Mathematical Reasoning
por: Oh, Minjae, et al.
Publicado: (2025)
por: Oh, Minjae, et al.
Publicado: (2025)
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
por: Choi, Yunho, et al.
Publicado: (2026)
por: Choi, Yunho, et al.
Publicado: (2026)
Where Should Diffusion Enter a Language Model? Geometry-Guided Hidden-State Replacement
por: Kong, Injin, et al.
Publicado: (2026)
por: Kong, Injin, et al.
Publicado: (2026)
Thinking Like a Doctor: Conversational Diagnosis through the Exploration of Diagnostic Knowledge Graphs
por: Won, Jeongmoon, et al.
Publicado: (2026)
por: Won, Jeongmoon, et al.
Publicado: (2026)
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
por: Lee, Yooseop, et al.
Publicado: (2025)
por: Lee, Yooseop, et al.
Publicado: (2025)
Quantifying Data Contamination in Psychometric Evaluations of LLMs
por: Han, Jongwook, et al.
Publicado: (2025)
por: Han, Jongwook, et al.
Publicado: (2025)
Mechanism Shift During Post-training from Autoregressive to Masked Diffusion Language Models
por: Kong, Injin, et al.
Publicado: (2026)
por: Kong, Injin, et al.
Publicado: (2026)
Psychometric Item Validation Using Virtual Respondents with Trait-Response Mediators
por: Lim, Sungjib, et al.
Publicado: (2025)
por: Lim, Sungjib, et al.
Publicado: (2025)
Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the Models
por: Lee, Jonggeun, et al.
Publicado: (2025)
por: Lee, Jonggeun, et al.
Publicado: (2025)
SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?
por: Lee, Jonggeun, et al.
Publicado: (2026)
por: Lee, Jonggeun, et al.
Publicado: (2026)
SpokenUS: A Spoken User Simulator for Task-Oriented Dialogue
por: Lee, Jonggeun, et al.
Publicado: (2026)
por: Lee, Jonggeun, et al.
Publicado: (2026)
Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
por: Park, Yoonah, et al.
Publicado: (2025)
por: Park, Yoonah, et al.
Publicado: (2025)
SUIT: Knowledge Editing with Subspace-Aware Key-Value Mappings
por: Park, Haewon, et al.
Publicado: (2025)
por: Park, Haewon, et al.
Publicado: (2025)
Stress-Testing Emotional Support Models: Moving from Homogeneous to Diverse Help Seekers
por: Heo, Chaewon, et al.
Publicado: (2026)
por: Heo, Chaewon, et al.
Publicado: (2026)
Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
por: Song, Jiwon, et al.
Publicado: (2025)
por: Song, Jiwon, et al.
Publicado: (2025)
Improving Dialogue State Tracking through Combinatorial Search for In-Context Examples
por: Pyun, Haesung, et al.
Publicado: (2025)
por: Pyun, Haesung, et al.
Publicado: (2025)
ToolDial: Multi-turn Dialogue Generation Method for Tool-Augmented Language Models
por: Shim, Jeonghoon, et al.
Publicado: (2025)
por: Shim, Jeonghoon, et al.
Publicado: (2025)
Dialogue Systems for Emotional Support via Value Reinforcement
por: Kim, Juhee, et al.
Publicado: (2025)
por: Kim, Juhee, et al.
Publicado: (2025)
Rewarding How Models Think Pedagogically: Integrating Pedagogical Reasoning and Thinking Rewards for LLMs in Education
por: Lee, Unggi, et al.
Publicado: (2026)
por: Lee, Unggi, et al.
Publicado: (2026)
Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid Items
por: Han, Jongwook, et al.
Publicado: (2025)
por: Han, Jongwook, et al.
Publicado: (2025)
Non-Collaborative User Simulators for Tool Agents
por: Shim, Jeonghoon, et al.
Publicado: (2025)
por: Shim, Jeonghoon, et al.
Publicado: (2025)
Model-based Preference Optimization in Abstractive Summarization without Human Feedback
por: Choi, Jaepill, et al.
Publicado: (2024)
por: Choi, Jaepill, et al.
Publicado: (2024)
Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval
por: Song, Jonghyun, et al.
Publicado: (2025)
por: Song, Jonghyun, et al.
Publicado: (2025)
Human Psychometric Questionnaires Mischaracterize LLM Behavior
por: Song, Woojung, et al.
Publicado: (2025)
por: Song, Woojung, et al.
Publicado: (2025)
Generalizing Visual Question Answering from Synthetic to Human-Written Questions via a Chain of QA with a Large Language Model
por: Kim, Taehee, et al.
Publicado: (2024)
por: Kim, Taehee, et al.
Publicado: (2024)
Context-Robust Knowledge Editing for Language Models
por: Park, Haewon, et al.
Publicado: (2025)
por: Park, Haewon, et al.
Publicado: (2025)
Mitigating Hallucination in Abstractive Summarization with Domain-Conditional Mutual Information
por: Chae, Kyubyung, et al.
Publicado: (2024)
por: Chae, Kyubyung, et al.
Publicado: (2024)
Learning to Retrieve User History and Generate User Profiles for Personalized Persuasiveness Prediction
por: Park, Sejun, et al.
Publicado: (2026)
por: Park, Sejun, et al.
Publicado: (2026)
Prompt Architecture Determines Reasoning Quality: A Variable Isolation Study on the Car Wash Problem
por: Jo, Heejin
Publicado: (2026)
por: Jo, Heejin
Publicado: (2026)
SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
por: Seo, Gyuhyeon, et al.
Publicado: (2025)
por: Seo, Gyuhyeon, et al.
Publicado: (2025)
Personalized LLM Decoding via Contrasting Personal Preference
por: Bu, Hyungjune, et al.
Publicado: (2025)
por: Bu, Hyungjune, et al.
Publicado: (2025)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
por: Jo, Dongwon, et al.
Publicado: (2025)
por: Jo, Dongwon, et al.
Publicado: (2025)
PVP: An Image Dataset for Personalized Visual Persuasion with Persuasion Strategies, Viewer Characteristics, and Persuasiveness Ratings
por: Kim, Junseo, et al.
Publicado: (2025)
por: Kim, Junseo, et al.
Publicado: (2025)
ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search
por: Lee, Hyunseok, et al.
Publicado: (2025)
por: Lee, Hyunseok, et al.
Publicado: (2025)
The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
por: Lee, Seongyun, et al.
Publicado: (2025)
por: Lee, Seongyun, et al.
Publicado: (2025)
A2SF: Accumulative Attention Scoring with Forgetting Factor for Token Pruning in Transformer Decoder
por: Jo, Hyun-rae, et al.
Publicado: (2024)
por: Jo, Hyun-rae, et al.
Publicado: (2024)
TABED: Test-Time Adaptive Ensemble Drafting for Robust Speculative Decoding in LVLMs
por: Lee, Minjae, et al.
Publicado: (2026)
por: Lee, Minjae, et al.
Publicado: (2026)
LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation
por: Park, Junyeong, et al.
Publicado: (2025)
por: Park, Junyeong, et al.
Publicado: (2025)
Ejemplares similares
-
KL for a KL: On-Policy Distillation with Control Variate Baseline
por: Oh, Minjae, et al.
Publicado: (2026) -
In-N-Out: A Parameter-Level API Graph Dataset for Tool Agents
por: Lee, Seungkyu, et al.
Publicado: (2025) -
Future Policy Approximation for Offline Reinforcement Learning Improves Mathematical Reasoning
por: Oh, Minjae, et al.
Publicado: (2025) -
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
por: Choi, Yunho, et al.
Publicado: (2026) -
Where Should Diffusion Enter a Language Model? Geometry-Guided Hidden-State Replacement
por: Kong, Injin, et al.
Publicado: (2026)