Rethinking Reward Models for Multi-Domain Test-Time Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Dong Bok, Lee, Seanie, Park, Sangwoo, Kang, Minki, Baek, Jinheon, Kim, Dongki, Wagner, Dominik, Jin, Jiongdao, Lee, Heejun, Bocklet, Tobias, Wang, Jinyu, Fu, Jingjing, Hwang, Sung Ju, Bian, Jiang, Song, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FedSVD: Adaptive Orthogonalization for Private Federated Learning with LoRA
by: Lee, Seanie, et al.
Published: (2025)
by: Lee, Seanie, et al.
Published: (2025)
SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models
by: Lee, Seanie, et al.
Published: (2025)
by: Lee, Seanie, et al.
Published: (2025)
OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources
by: Baek, Jinheon, et al.
Published: (2026)
by: Baek, Jinheon, et al.
Published: (2026)
PREPING: Building Agent Memory without Tasks
by: Choi, Yumin, et al.
Published: (2026)
by: Choi, Yumin, et al.
Published: (2026)
FedRand: Enhancing Privacy in Federated Learning with Randomized LoRA Subparameter Updates
by: Park, Sangwoo, et al.
Published: (2025)
by: Park, Sangwoo, et al.
Published: (2025)
Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs
by: Choi, Yumin, et al.
Published: (2025)
by: Choi, Yumin, et al.
Published: (2025)
HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
by: Lee, Seanie, et al.
Published: (2024)
by: Lee, Seanie, et al.
Published: (2024)
Chain of Retrieval: Multi-Aspect Iterative Search Expansion and Post-Order Search Aggregation for Full Paper Retrieval
by: Park, Sangwoo, et al.
Published: (2025)
by: Park, Sangwoo, et al.
Published: (2025)
It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs
by: Park, Sangwoo, et al.
Published: (2026)
by: Park, Sangwoo, et al.
Published: (2026)
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR
by: Lee, Chanuk, et al.
Published: (2026)
by: Lee, Chanuk, et al.
Published: (2026)
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
by: Lee, Hyomin, et al.
Published: (2026)
by: Lee, Hyomin, et al.
Published: (2026)
Rethinking Code Refinement: Learning to Judge Code Efficiency
by: Seo, Minju, et al.
Published: (2024)
by: Seo, Minju, et al.
Published: (2024)
Distilling LLM Agent into Small Models with Retrieval and Code Tools
by: Kang, Minki, et al.
Published: (2025)
by: Kang, Minki, et al.
Published: (2025)
Optimized Speculative Sampling for GPU Hardware Accelerators
by: Wagner, Dominik, et al.
Published: (2024)
by: Wagner, Dominik, et al.
Published: (2024)
THINKSAFE: Self-Generated Safety Alignment for Reasoning Models
by: Lee, Seanie, et al.
Published: (2026)
by: Lee, Seanie, et al.
Published: (2026)
Efficient Long Context Language Model Retrieval with Compression
by: Seo, Minju, et al.
Published: (2024)
by: Seo, Minju, et al.
Published: (2024)
Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning
by: Seo, Minju, et al.
Published: (2025)
by: Seo, Minju, et al.
Published: (2025)
Self-Supervised Dataset Distillation for Transfer Learning
by: Lee, Dong Bok, et al.
Published: (2023)
by: Lee, Dong Bok, et al.
Published: (2023)
Mol-LLaMA: Towards General Understanding of Molecules in Large Molecular Language Model
by: Kim, Dongki, et al.
Published: (2025)
by: Kim, Dongki, et al.
Published: (2025)
ES-Merging: Biological MLLM Merging via Embedding Space Signals
by: Lee, Wonbin, et al.
Published: (2026)
by: Lee, Wonbin, et al.
Published: (2026)
Retrieval-Augmented Data Augmentation for Low-Resource Domain Tasks
by: Seo, Minju, et al.
Published: (2024)
by: Seo, Minju, et al.
Published: (2024)
Drug Discovery with Dynamic Goal-aware Fragments
by: Lee, Seul, et al.
Published: (2023)
by: Lee, Seul, et al.
Published: (2023)
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
by: Wagner, Dominik, et al.
Published: (2025)
by: Wagner, Dominik, et al.
Published: (2025)
Efficient Real-time Refinement of Language Model Text Generation
by: Ko, Joonho, et al.
Published: (2025)
by: Ko, Joonho, et al.
Published: (2025)
System Prompt Optimization with Meta-Learning
by: Choi, Yumin, et al.
Published: (2025)
by: Choi, Yumin, et al.
Published: (2025)
Unified Multimodal Interleaved Document Representation for Retrieval
by: Lee, Jaewoo, et al.
Published: (2024)
by: Lee, Jaewoo, et al.
Published: (2024)
Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction
by: Willette, Jeffrey, et al.
Published: (2025)
by: Willette, Jeffrey, et al.
Published: (2025)
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU
by: Lee, Heejun, et al.
Published: (2025)
by: Lee, Heejun, et al.
Published: (2025)
DiffusionNAG: Predictor-guided Neural Architecture Generation with Diffusion Models
by: An, Sohyun, et al.
Published: (2023)
by: An, Sohyun, et al.
Published: (2023)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
by: Lee, Chanuk, et al.
Published: (2026)
by: Lee, Chanuk, et al.
Published: (2026)
Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching
by: Aytes, Simon A., et al.
Published: (2025)
by: Aytes, Simon A., et al.
Published: (2025)
Set-based Meta-Interpolation for Few-Task Meta-Learning
by: Lee, Seanie, et al.
Published: (2022)
by: Lee, Seanie, et al.
Published: (2022)
GFlowPO: Generative Flow Network as a Language Model Prompt Optimizer
by: Cho, Junmo, et al.
Published: (2026)
by: Cho, Junmo, et al.
Published: (2026)
VideoRAG: Retrieval-Augmented Generation over Video Corpus
by: Jeong, Soyeong, et al.
Published: (2025)
by: Jeong, Soyeong, et al.
Published: (2025)
SEA: Sparse Linear Attention with Estimated Attention Mask
by: Lee, Heejun, et al.
Published: (2023)
by: Lee, Heejun, et al.
Published: (2023)
Training-Free Exponential Context Extension via Cascading KV Cache
by: Willette, Jeffrey, et al.
Published: (2024)
by: Willette, Jeffrey, et al.
Published: (2024)
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
by: Lee, Youngwan, et al.
Published: (2025)
by: Lee, Youngwan, et al.
Published: (2025)
Latent Paraphrasing: Perturbation on Layers Improves Knowledge Injection in Language Models
by: Kang, Minki, et al.
Published: (2024)
by: Kang, Minki, et al.
Published: (2024)
Database-Augmented Query Representation for Information Retrieval
by: Jeong, Soyeong, et al.
Published: (2024)
by: Jeong, Soyeong, et al.
Published: (2024)
Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity
by: Jeong, Soyeong, et al.
Published: (2024)
by: Jeong, Soyeong, et al.
Published: (2024)
Similar Items
-
FedSVD: Adaptive Orthogonalization for Private Federated Learning with LoRA
by: Lee, Seanie, et al.
Published: (2025) -
SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models
by: Lee, Seanie, et al.
Published: (2025) -
OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources
by: Baek, Jinheon, et al.
Published: (2026) -
PREPING: Building Agent Memory without Tasks
by: Choi, Yumin, et al.
Published: (2026) -
FedRand: Enhancing Privacy in Federated Learning with Randomized LoRA Subparameter Updates
by: Park, Sangwoo, et al.
Published: (2025)