Gnothi Seauton: Empowering Faithful Self-Interpretability in Black-Box Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Shaobo, Tang, Hongxuan, Wang, Mingyang, Zhang, Hongrui, Liu, Xuyang, Li, Weiya, Hu, Xuming, Zhang, Linfeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs
von: Zhou, Yufa, et al.
Veröffentlicht: (2025)
von: Zhou, Yufa, et al.
Veröffentlicht: (2025)
Online Learning and Equilibrium Computation with Ranking Feedback
von: Liu, Mingyang, et al.
Veröffentlicht: (2026)
von: Liu, Mingyang, et al.
Veröffentlicht: (2026)
Measuring Bargaining Abilities of LLMs: A Benchmark and A Buyer-Enhancement Method
von: Xia, Tian, et al.
Veröffentlicht: (2024)
von: Xia, Tian, et al.
Veröffentlicht: (2024)
Peer-Predictive Self-Training for Language Model Reasoning
von: Feng, Shi, et al.
Veröffentlicht: (2026)
von: Feng, Shi, et al.
Veröffentlicht: (2026)
Fairshare Data Pricing via Data Valuation for Large Language Models
von: Zhang, Luyang, et al.
Veröffentlicht: (2025)
von: Zhang, Luyang, et al.
Veröffentlicht: (2025)
Beyond Black-Box Interventions: Latent Probing for Faithful Retrieval-Augmented Generation
von: Gao, Linfeng, et al.
Veröffentlicht: (2025)
von: Gao, Linfeng, et al.
Veröffentlicht: (2025)
Make an Offer They Can't Refuse: Grounding Bayesian Persuasion in Real-World Dialogues without Pre-Commitment
von: He, Buwei, et al.
Veröffentlicht: (2025)
von: He, Buwei, et al.
Veröffentlicht: (2025)
Incentivizing Inclusive Contributions in Model Sharing Markets
von: Zhang, Enpei, et al.
Veröffentlicht: (2025)
von: Zhang, Enpei, et al.
Veröffentlicht: (2025)
ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models
von: Chen, Junzhe, et al.
Veröffentlicht: (2024)
von: Chen, Junzhe, et al.
Veröffentlicht: (2024)
Black-Box Lifting and Robustness Theorems for Multi-Agent Contracts
von: Dütting, Paul, et al.
Veröffentlicht: (2025)
von: Dütting, Paul, et al.
Veröffentlicht: (2025)
ShortageSim: Simulating Drug Shortages under Information Asymmetry
von: Cui, Mingxuan, et al.
Veröffentlicht: (2025)
von: Cui, Mingxuan, et al.
Veröffentlicht: (2025)
Language Self-Play For Data-Free Training
von: Kuba, Jakub Grudzien, et al.
Veröffentlicht: (2025)
von: Kuba, Jakub Grudzien, et al.
Veröffentlicht: (2025)
Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024)
Nash CoT: Multi-Path Inference with Preference Equilibrium
von: Zhang, Ziqi, et al.
Veröffentlicht: (2024)
von: Zhang, Ziqi, et al.
Veröffentlicht: (2024)
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
Shifting AI Efficiency From Model-Centric to Data-Centric Compression
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
Competitive Information Design for Pandora's Box
von: Ding, Bolin, et al.
Veröffentlicht: (2021)
von: Ding, Bolin, et al.
Veröffentlicht: (2021)
UGen: Unified Autoregressive Multimodal Model with Progressive Vocabulary Learning
von: Tang, Hongxuan, et al.
Veröffentlicht: (2025)
von: Tang, Hongxuan, et al.
Veröffentlicht: (2025)
STRIDE: A Tool-Assisted LLM Agent Framework for Strategic and Interactive Decision-Making
von: Li, Chuanhao, et al.
Veröffentlicht: (2024)
von: Li, Chuanhao, et al.
Veröffentlicht: (2024)
Beyond Scalar Rewards: Dense Feedback for LLM Policy Synthesis in Sequential Social Dilemmas
von: Gallego, Víctor
Veröffentlicht: (2026)
von: Gallego, Víctor
Veröffentlicht: (2026)
Evolving Diverse Red-team Language Models in Multi-round Multi-agent Games
von: Ma, Chengdong, et al.
Veröffentlicht: (2023)
von: Ma, Chengdong, et al.
Veröffentlicht: (2023)
Verification Required: The Impact of Information Credibility on AI Persuasion
von: Mahmud, Saaduddin, et al.
Veröffentlicht: (2026)
von: Mahmud, Saaduddin, et al.
Veröffentlicht: (2026)
Emergent LLM behaviors are observationally equivalent to data leakage
von: Barrie, Christopher, et al.
Veröffentlicht: (2025)
von: Barrie, Christopher, et al.
Veröffentlicht: (2025)
Beyond Game Theory Optimal: Profit-Maximizing Poker Agents for No-Limit Holdem
von: Yi, SeungHyun, et al.
Veröffentlicht: (2025)
von: Yi, SeungHyun, et al.
Veröffentlicht: (2025)
Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
ALYMPICS: LLM Agents Meet Game Theory -- Exploring Strategic Decision-Making with AI Agents
von: Mao, Shaoguang, et al.
Veröffentlicht: (2023)
von: Mao, Shaoguang, et al.
Veröffentlicht: (2023)
Black-Box Followers, White-Box Leaders: Partial Zeroth-Order Methods for MPECs
von: Fischer, Miriam, et al.
Veröffentlicht: (2026)
von: Fischer, Miriam, et al.
Veröffentlicht: (2026)
Eliciting Informative Text Evaluations with Large Language Models
von: Lu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Lu, Yuxuan, et al.
Veröffentlicht: (2024)
Mediator Interpretation and Faster Learning Algorithms for Linear Correlated Equilibria in General Extensive-Form Games
von: Zhang, Brian Hu, et al.
Veröffentlicht: (2023)
von: Zhang, Brian Hu, et al.
Veröffentlicht: (2023)
BEDA: Belief Estimation as Probabilistic Constraints for Performing Strategic Dialogue Acts
von: Li, Hengli, et al.
Veröffentlicht: (2025)
von: Li, Hengli, et al.
Veröffentlicht: (2025)
The Value of Context: Human versus Black Box Evaluators
von: Iakovlev, Andrei, et al.
Veröffentlicht: (2024)
von: Iakovlev, Andrei, et al.
Veröffentlicht: (2024)
AI for Service: Proactive Assistance with AI Glasses
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
Improving Community-Participated Patrol for Anti-Poaching
von: Wu, Yufei, et al.
Veröffentlicht: (2024)
von: Wu, Yufei, et al.
Veröffentlicht: (2024)
MGTUNet: An new UNet for colon nuclei instance segmentation and quantification
von: Pan, Liangrui, et al.
Veröffentlicht: (2022)
von: Pan, Liangrui, et al.
Veröffentlicht: (2022)
Multi-Fidelity Bayesian Optimization for Nash Equilibria with Black-Box Utilities
von: Zhang, Yunchuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yunchuan, et al.
Veröffentlicht: (2025)
Price Stability and Improved Buyer Utility with Presentation Design: A Theoretical Study of the Amazon Buy Box
von: Friedler, Ophir, et al.
Veröffentlicht: (2025)
von: Friedler, Ophir, et al.
Veröffentlicht: (2025)
Private Private Information in Second-Price Auction
von: Liu, Boyu, et al.
Veröffentlicht: (2026)
von: Liu, Boyu, et al.
Veröffentlicht: (2026)
A Transformer-Based Neural Network for Optimal Deterministic-Allocation and Anonymous Joint Auction Design
von: Zhang, Zhen, et al.
Veröffentlicht: (2025)
von: Zhang, Zhen, et al.
Veröffentlicht: (2025)
From Natural Language to Extensive-Form Game Representations
von: Deng, Shilong, et al.
Veröffentlicht: (2025)
von: Deng, Shilong, et al.
Veröffentlicht: (2025)
Incentive-Aligned Multi-Source LLM Summaries
von: Jiang, Yanchen, et al.
Veröffentlicht: (2025)
von: Jiang, Yanchen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs
von: Zhou, Yufa, et al.
Veröffentlicht: (2025) -
Online Learning and Equilibrium Computation with Ranking Feedback
von: Liu, Mingyang, et al.
Veröffentlicht: (2026) -
Measuring Bargaining Abilities of LLMs: A Benchmark and A Buyer-Enhancement Method
von: Xia, Tian, et al.
Veröffentlicht: (2024) -
Peer-Predictive Self-Training for Language Model Reasoning
von: Feng, Shi, et al.
Veröffentlicht: (2026) -
Fairshare Data Pricing via Data Valuation for Large Language Models
von: Zhang, Luyang, et al.
Veröffentlicht: (2025)