From Actions to Words: Towards Abstractive-Textual Policy Summarization in RL
Fuente:
arXiv
Saved in:
| Main Authors: | Admoni, Sahar, Hallak, Assaf, Ziser, Yftah, Ben-Porat, Omer, Amir, Ofra |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aligning What LLMs Do and Say: Towards Self-Consistent Explanations
by: Admoni, Sahar, et al.
Published: (2025)
by: Admoni, Sahar, et al.
Published: (2025)
More Test-Time Compute Can Hurt: Overestimation Bias in LLM Beam Search
by: Dalal, Gal, et al.
Published: (2026)
by: Dalal, Gal, et al.
Published: (2026)
From Competition to Collaboration: Designing Sustainable Mechanisms Between LLMs and Online Forums
by: Fono, Niv, et al.
Published: (2026)
by: Fono, Niv, et al.
Published: (2026)
Assessing Policy Updates: Toward Trust-Preserving Intelligent User Interfaces
by: Solomon, Matan, et al.
Published: (2025)
by: Solomon, Matan, et al.
Published: (2025)
When Knowledge Is Not Free: Cost-Aware Evidence Selection in Retrieval-Augmented Generation
by: Wu, Mingyan, et al.
Published: (2026)
by: Wu, Mingyan, et al.
Published: (2026)
Personalized Reinforcement Learning with a Budget of Policies
by: Ivanov, Dmitry, et al.
Published: (2024)
by: Ivanov, Dmitry, et al.
Published: (2024)
Layer by Layer: Uncovering Where Multi-Task Learning Happens in Instruction-Tuned Large Language Models
by: Zhao, Zheng, et al.
Published: (2024)
by: Zhao, Zheng, et al.
Published: (2024)
SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs
by: Ghosh, Shaona, et al.
Published: (2025)
by: Ghosh, Shaona, et al.
Published: (2025)
"Trust me on this" Explaining Agent Behavior to a Human Terminator
by: Menkes, Uri, et al.
Published: (2025)
by: Menkes, Uri, et al.
Published: (2025)
INSIGHTS: Demonstration-Based Summaries of Time Series Predictors
by: Porat, Bar Eini, et al.
Published: (2026)
by: Porat, Bar Eini, et al.
Published: (2026)
Neural Message-Passing on Attention Graphs for Hallucination Detection
by: Frasca, Fabrizio, et al.
Published: (2025)
by: Frasca, Fabrizio, et al.
Published: (2025)
RL-RC-DoT: A Block-level RL agent for Task-Aware Video Compression
by: Gadot, Uri, et al.
Published: (2025)
by: Gadot, Uri, et al.
Published: (2025)
Who Said Neural Networks Aren't Linear?
by: Berman, Nimrod, et al.
Published: (2025)
by: Berman, Nimrod, et al.
Published: (2025)
Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
by: Bar-Shalom, Guy, et al.
Published: (2025)
by: Bar-Shalom, Guy, et al.
Published: (2025)
Modeling Churn in Recommender Systems with Aggregated Preferences
by: Keinan, Gur, et al.
Published: (2025)
by: Keinan, Gur, et al.
Published: (2025)
Self-Improving World Modelling with Latent Actions
by: Qiu, Yifu, et al.
Published: (2026)
by: Qiu, Yifu, et al.
Published: (2026)
Policy Gradient with Tree Expansion
by: Dalal, Gal, et al.
Published: (2023)
by: Dalal, Gal, et al.
Published: (2023)
BRIDO: Bringing Democratic Order to Abstractive Summarization
by: Lee, Junhyun, et al.
Published: (2025)
by: Lee, Junhyun, et al.
Published: (2025)
Bandits with Single-Peaked Preferences and Limited Resources
by: Ben-Porat, Omer, et al.
Published: (2025)
by: Ben-Porat, Omer, et al.
Published: (2025)
Envious Explore and Exploit
by: Ben-Porat, Omer, et al.
Published: (2025)
by: Ben-Porat, Omer, et al.
Published: (2025)
Integrating Policy Summaries with Reward Decomposition for Explaining Reinforcement Learning Agents
by: Septon, Yael, et al.
Published: (2022)
by: Septon, Yael, et al.
Published: (2022)
LOCOST: State-Space Models for Long Document Abstractive Summarization
by: Bronnec, Florian Le, et al.
Published: (2024)
by: Bronnec, Florian Le, et al.
Published: (2024)
Abstractive Text Summarization: State of the Art, Challenges, and Improvements
by: Shakil, Hassan, et al.
Published: (2024)
by: Shakil, Hassan, et al.
Published: (2024)
Spectral Editing of Activations for Large Language Model Alignment
by: Qiu, Yifu, et al.
Published: (2024)
by: Qiu, Yifu, et al.
Published: (2024)
Incorporating Distributions of Discourse Structure for Long Document Abstractive Summarization
by: Liu, Dongqi, et al.
Published: (2023)
by: Liu, Dongqi, et al.
Published: (2023)
Beyond Next Token Probabilities: Learnable, Fast Detection of Hallucinations and Data Contamination on LLM Output Distributions
by: Bar-Shalom, Guy, et al.
Published: (2025)
by: Bar-Shalom, Guy, et al.
Published: (2025)
Conformalized Teleoperation: Confidently Mapping Human Inputs to High-Dimensional Robot Actions
by: Zhao, Michelle, et al.
Published: (2024)
by: Zhao, Michelle, et al.
Published: (2024)
Improving Faithfulness of Abstractive Summarization by Controlling Confounding Effect of Irrelevant Sentences
by: Ghoshal, Asish, et al.
Published: (2022)
by: Ghoshal, Asish, et al.
Published: (2022)
Salient Information Prompting to Steer Content in Prompt-based Abstractive Summarization
by: Xu, Lei, et al.
Published: (2024)
by: Xu, Lei, et al.
Published: (2024)
Adapting by Analogy: OOD Generalization of Visuomotor Policies via Functional Correspondence
by: Gupta, Pranay, et al.
Published: (2025)
by: Gupta, Pranay, et al.
Published: (2025)
Entity-level Factual Adaptiveness of Fine-tuning based Abstractive Summarization Models
by: Song, Jongyoon, et al.
Published: (2024)
by: Song, Jongyoon, et al.
Published: (2024)
uMedSum: A Unified Framework for Advancing Medical Abstractive Summarization
by: Nagar, Aishik, et al.
Published: (2024)
by: Nagar, Aishik, et al.
Published: (2024)
STATe-of-Thoughts: Structured Action Templates for Tree-of-Thoughts
by: Bamberger, Zachary, et al.
Published: (2026)
by: Bamberger, Zachary, et al.
Published: (2026)
Modeling Attrition in Recommender Systems with Departing Bandits
by: Ben-Porat, Omer, et al.
Published: (2022)
by: Ben-Porat, Omer, et al.
Published: (2022)
A Second Look on BASS -- Boosting Abstractive Summarization with Unified Semantic Graphs -- A Replication Study
by: Koraş, Osman Alperen, et al.
Published: (2024)
by: Koraş, Osman Alperen, et al.
Published: (2024)
RL with Learnable Textual Feedback: A Bilevel Approach
by: Singh, Utsav, et al.
Published: (2026)
by: Singh, Utsav, et al.
Published: (2026)
Uncovering RL Integration in SSL Loss: Objective-Specific Implications for Data-Efficient RL
by: Çağatan, Ömer Veysel, et al.
Published: (2024)
by: Çağatan, Ömer Veysel, et al.
Published: (2024)
A Comprehensive Machine Learning Framework for Micromobility Demand Prediction
by: Porat, Omri, et al.
Published: (2025)
by: Porat, Omri, et al.
Published: (2025)
Improving Sequence-to-Sequence Models for Abstractive Text Summarization Using Meta Heuristic Approaches
by: Saxena, Aditya, et al.
Published: (2024)
by: Saxena, Aditya, et al.
Published: (2024)
L3Cube-MahaSum: A Comprehensive Dataset and BART Models for Abstractive Text Summarization in Marathi
by: Deshmukh, Pranita, et al.
Published: (2024)
by: Deshmukh, Pranita, et al.
Published: (2024)
Similar Items
-
Aligning What LLMs Do and Say: Towards Self-Consistent Explanations
by: Admoni, Sahar, et al.
Published: (2025) -
More Test-Time Compute Can Hurt: Overestimation Bias in LLM Beam Search
by: Dalal, Gal, et al.
Published: (2026) -
From Competition to Collaboration: Designing Sustainable Mechanisms Between LLMs and Online Forums
by: Fono, Niv, et al.
Published: (2026) -
Assessing Policy Updates: Toward Trust-Preserving Intelligent User Interfaces
by: Solomon, Matan, et al.
Published: (2025) -
When Knowledge Is Not Free: Cost-Aware Evidence Selection in Retrieval-Augmented Generation
by: Wu, Mingyan, et al.
Published: (2026)