Guardado en:
| Autor principal: | Xu, Jiexi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2509.25267 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents
por: Chua, Jaymari, et al.
Publicado: (2025)
por: Chua, Jaymari, et al.
Publicado: (2025)
Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy
por: He, Langzhou, et al.
Publicado: (2026)
por: He, Langzhou, et al.
Publicado: (2026)
Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation
por: Koc, Vincent
Publicado: (2025)
por: Koc, Vincent
Publicado: (2025)
GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
por: Agrawal, Lakshya A, et al.
Publicado: (2025)
por: Agrawal, Lakshya A, et al.
Publicado: (2025)
Survey Transfer Learning: Recycling Data with Silicon Responses
por: Amini, Ali
Publicado: (2025)
por: Amini, Ali
Publicado: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
por: Estevanell-Valladares, Ernesto L., et al.
Publicado: (2025)
por: Estevanell-Valladares, Ernesto L., et al.
Publicado: (2025)
Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
por: Nakamura, Mason, et al.
Publicado: (2025)
por: Nakamura, Mason, et al.
Publicado: (2025)
ReaGeo: Reasoning-Enhanced End-to-End Geocoding with LLMs
por: Cui, Jian, et al.
Publicado: (2026)
por: Cui, Jian, et al.
Publicado: (2026)
SCULPT: Constraint-Guided Pruned MCTS that Carves Efficient Paths for Mathematical Reasoning
por: Fang, Qitong, et al.
Publicado: (2026)
por: Fang, Qitong, et al.
Publicado: (2026)
On the Limits of Learned Importance Scoring for KV Cache Compression
por: Steele, Brady
Publicado: (2026)
por: Steele, Brady
Publicado: (2026)
ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
por: Zhou, Yue, et al.
Publicado: (2025)
por: Zhou, Yue, et al.
Publicado: (2025)
Scaling Trends for Multi-Hop Contextual Reasoning in Mid-Scale Language Models
por: Steele, Brady, et al.
Publicado: (2026)
por: Steele, Brady, et al.
Publicado: (2026)
The Reasoning-Creativity Trade-off: Toward Creativity-Driven Problem Solving
por: Luyten, Max Ruiz, et al.
Publicado: (2026)
por: Luyten, Max Ruiz, et al.
Publicado: (2026)
Learning, Fast and Slow: Towards LLMs That Adapt Continually
por: Tiwari, Rishabh, et al.
Publicado: (2026)
por: Tiwari, Rishabh, et al.
Publicado: (2026)
When Gradients Collide: Failure Modes of Multi-Objective Prompt Optimization for LLM Judges
por: Darshan, Parth, et al.
Publicado: (2026)
por: Darshan, Parth, et al.
Publicado: (2026)
Evo-DKD: Dual-Knowledge Decoding for Autonomous Ontology Evolution in Large Language Models
por: Raman, Vishal, et al.
Publicado: (2025)
por: Raman, Vishal, et al.
Publicado: (2025)
Deployment-Time Reliability of Learned Robot Policies
por: Agia, Christopher
Publicado: (2026)
por: Agia, Christopher
Publicado: (2026)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
por: Yousaf, Iqra
Publicado: (2024)
por: Yousaf, Iqra
Publicado: (2024)
GraphEval36K: Benchmarking Coding and Reasoning Capabilities of Large Language Models on Graph Datasets
por: Wu, Qiming, et al.
Publicado: (2024)
por: Wu, Qiming, et al.
Publicado: (2024)
MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization
por: Tanjim, Md Mehrab, et al.
Publicado: (2026)
por: Tanjim, Md Mehrab, et al.
Publicado: (2026)
Dynamics of COVID-19 Misinformation: An Analysis of Conspiracy Theories, Fake Remedies, and False Reports
por: Thakur, Nirmalya, et al.
Publicado: (2025)
por: Thakur, Nirmalya, et al.
Publicado: (2025)
Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture
por: Iscan, Mehmet
Publicado: (2026)
por: Iscan, Mehmet
Publicado: (2026)
The Single-File Test: A Longitudinal Public-Interface Evaluation of First-Output LLM Web Generation with Social Reach Tracking
por: Palacios, Diego Cabezas
Publicado: (2026)
por: Palacios, Diego Cabezas
Publicado: (2026)
Bridging the Gap: An Intermediate Language for Enhanced and Cost-Effective Grapheme-to-Phoneme Conversion with Homographs with Multiple Pronunciations Disambiguation
por: Bertina, Abbas, et al.
Publicado: (2025)
por: Bertina, Abbas, et al.
Publicado: (2025)
Streaming Continual Learning for Unified Adaptive Intelligence in Dynamic Environments
por: Giannini, Federico, et al.
Publicado: (2026)
por: Giannini, Federico, et al.
Publicado: (2026)
Procedural Game Level Design with Deep Reinforcement Learning
por: Özkan, Miraç Buğra
Publicado: (2025)
por: Özkan, Miraç Buğra
Publicado: (2025)
Predictive Analytics for Collaborators Answers, Code Quality, and Dropout on Stack Overflow
por: Zolduoarrati, Elijah, et al.
Publicado: (2025)
por: Zolduoarrati, Elijah, et al.
Publicado: (2025)
An Aircraft Upset Recovery System with Reinforcement Learning
por: Demir, Mahir, et al.
Publicado: (2026)
por: Demir, Mahir, et al.
Publicado: (2026)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
por: Pather, Kaviraj, et al.
Publicado: (2025)
por: Pather, Kaviraj, et al.
Publicado: (2025)
Bridging the Reasoning Gap: Small LLMs Can Plan with Generalised Strategies
por: Borro, Andrey, et al.
Publicado: (2025)
por: Borro, Andrey, et al.
Publicado: (2025)
Unveiling Hidden Factors: Explainable AI for Feature Boosting in Speech Emotion Recognition
por: Nfissi, Alaa, et al.
Publicado: (2024)
por: Nfissi, Alaa, et al.
Publicado: (2024)
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
por: Radosky, Lukas, et al.
Publicado: (2026)
por: Radosky, Lukas, et al.
Publicado: (2026)
Iterative Feature Boosting for Explainable Speech Emotion Recognition
por: Nfissi, Alaa, et al.
Publicado: (2024)
por: Nfissi, Alaa, et al.
Publicado: (2024)
Pioneer Agent: Continual Improvement of Small Language Models in Production
por: Atreja, Dhruv, et al.
Publicado: (2026)
por: Atreja, Dhruv, et al.
Publicado: (2026)
Primary Care Diagnoses as a Reliable Predictor for Orthopedic Surgical Interventions
por: Verma, Khushboo, et al.
Publicado: (2025)
por: Verma, Khushboo, et al.
Publicado: (2025)
Gyan: An Explainable Neuro-Symbolic Language Model
por: Srinivasan, Venkat, et al.
Publicado: (2026)
por: Srinivasan, Venkat, et al.
Publicado: (2026)
CoupleEvo: Evolving Heuristics for Coupled Optimization Problems Using Large Language Models
por: Bömer, Thomas, et al.
Publicado: (2026)
por: Bömer, Thomas, et al.
Publicado: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025)
por: Fadli, Samih
Publicado: (2025)
Retrieval Augmented Thought Process for Private Data Handling in Healthcare
por: Pouplin, Thomas, et al.
Publicado: (2024)
por: Pouplin, Thomas, et al.
Publicado: (2024)
Automated Circuit Interpretation via Probe Prompting
por: Birardi, Giuseppe
Publicado: (2025)
por: Birardi, Giuseppe
Publicado: (2025)
Ejemplares similares
-
Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents
por: Chua, Jaymari, et al.
Publicado: (2025) -
Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy
por: He, Langzhou, et al.
Publicado: (2026) -
Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation
por: Koc, Vincent
Publicado: (2025) -
GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
por: Agrawal, Lakshya A, et al.
Publicado: (2025) -
Survey Transfer Learning: Recycling Data with Silicon Responses
por: Amini, Ali
Publicado: (2025)