Gespeichert in:
| 1. Verfasser: | Lin, Hender |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2503.06648 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation
von: Bhattacharjee, Amrita, et al.
Veröffentlicht: (2024)
von: Bhattacharjee, Amrita, et al.
Veröffentlicht: (2024)
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
Context-Enhanced Contrastive Search for Improved LLM Text Generation
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
NLP Verification: Towards a General Methodology for Certifying Robustness
von: Casadio, Marco, et al.
Veröffentlicht: (2024)
von: Casadio, Marco, et al.
Veröffentlicht: (2024)
Muon is Scalable for LLM Training
von: Liu, Jingyuan, et al.
Veröffentlicht: (2025)
von: Liu, Jingyuan, et al.
Veröffentlicht: (2025)
LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
Bridging Robustness and Generalization Against Word Substitution Attacks in NLP via the Growth Bound Matrix Approach
von: Bouri, Mohammed, et al.
Veröffentlicht: (2025)
von: Bouri, Mohammed, et al.
Veröffentlicht: (2025)
CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment
von: Xie, Guofu, et al.
Veröffentlicht: (2025)
von: Xie, Guofu, et al.
Veröffentlicht: (2025)
EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models
von: Dhaini, Mahdi, et al.
Veröffentlicht: (2025)
von: Dhaini, Mahdi, et al.
Veröffentlicht: (2025)
RobustSentEmbed: Robust Sentence Embeddings Using Adversarial Self-Supervised Contrastive Learning
von: Asl, Javad Rafiei, et al.
Veröffentlicht: (2024)
von: Asl, Javad Rafiei, et al.
Veröffentlicht: (2024)
$C^2$: Scalable Auto-Feedback for LLM-based Chart Generation
von: Koh, Woosung, et al.
Veröffentlicht: (2024)
von: Koh, Woosung, et al.
Veröffentlicht: (2024)
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models Using Contrast Sets: An Experimental Approach
von: Sanwal, Manish
Veröffentlicht: (2024)
von: Sanwal, Manish
Veröffentlicht: (2024)
Automated Literature Review Using NLP Techniques and LLM-Based Retrieval-Augmented Generation
von: Ali, Nurshat Fateh, et al.
Veröffentlicht: (2024)
von: Ali, Nurshat Fateh, et al.
Veröffentlicht: (2024)
Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts
von: Yin, Yueqin, et al.
Veröffentlicht: (2024)
von: Yin, Yueqin, et al.
Veröffentlicht: (2024)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
von: Labiad, Ismail, et al.
Veröffentlicht: (2025)
von: Labiad, Ismail, et al.
Veröffentlicht: (2025)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
Robustly Improving LLM Fairness in Realistic Settings via Interpretability
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
Enhancing Annotated Bibliography Generation with LLM Ensembles
von: Bermejo, Sergio
Veröffentlicht: (2024)
von: Bermejo, Sergio
Veröffentlicht: (2024)
Post-training an LLM for RAG? Train on Self-Generated Demonstrations
von: Finlayson, Matthew, et al.
Veröffentlicht: (2025)
von: Finlayson, Matthew, et al.
Veröffentlicht: (2025)
ToxiGAN: Toxic Data Augmentation via LLM-Guided Directional Adversarial Generation
von: Li, Peiran, et al.
Veröffentlicht: (2026)
von: Li, Peiran, et al.
Veröffentlicht: (2026)
GEAR: A General Evaluation Framework for Abductive Reasoning
von: He, Kaiyu, et al.
Veröffentlicht: (2025)
von: He, Kaiyu, et al.
Veröffentlicht: (2025)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
von: Cui, Sijia, et al.
Veröffentlicht: (2026)
von: Cui, Sijia, et al.
Veröffentlicht: (2026)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
Training and Evaluating Language Models with Template-based Data Generation
von: Zhang, Yifan
Veröffentlicht: (2024)
von: Zhang, Yifan
Veröffentlicht: (2024)
Set-LLM: A Permutation-Invariant LLM
von: Egressy, Beni, et al.
Veröffentlicht: (2025)
von: Egressy, Beni, et al.
Veröffentlicht: (2025)
Federated Learning with Layer Skipping: Efficient Training of Large Language Models for Healthcare NLP
von: Zhang, Lihong, et al.
Veröffentlicht: (2025)
von: Zhang, Lihong, et al.
Veröffentlicht: (2025)
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
von: Karagoz, Atahan
Veröffentlicht: (2026)
von: Karagoz, Atahan
Veröffentlicht: (2026)
Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game
von: Cheng, Pengyu, et al.
Veröffentlicht: (2023)
von: Cheng, Pengyu, et al.
Veröffentlicht: (2023)
IPAD: Inverse Prompt for AI Detection - A Robust and Interpretable LLM-Generated Text Detector
von: Chen, Zheng, et al.
Veröffentlicht: (2025)
von: Chen, Zheng, et al.
Veröffentlicht: (2025)
Enhancing LLM Evaluations: The Garbling Trick
von: Bradley, William F.
Veröffentlicht: (2024)
von: Bradley, William F.
Veröffentlicht: (2024)
Interpretable AI for Time-Series: Multi-Model Heatmap Fusion with Global Attention and NLP-Generated Explanations
von: Francis, Jiztom Kavalakkatt, et al.
Veröffentlicht: (2025)
von: Francis, Jiztom Kavalakkatt, et al.
Veröffentlicht: (2025)
Adversarial Lens: Exploiting Attention Layers to Generate Adversarial Examples for Evaluation
von: Dhole, Kaustubh
Veröffentlicht: (2025)
von: Dhole, Kaustubh
Veröffentlicht: (2025)
SyGra: A Unified Graph-Based Framework for Scalable Generation, Quality Tagging, and Management of Synthetic Data
von: Pradhan, Bidyapati, et al.
Veröffentlicht: (2025)
von: Pradhan, Bidyapati, et al.
Veröffentlicht: (2025)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
CodeRefine: A Pipeline for Enhancing LLM-Generated Code Implementations of Research Papers
von: Trofimova, Ekaterina, et al.
Veröffentlicht: (2024)
von: Trofimova, Ekaterina, et al.
Veröffentlicht: (2024)
From Text to Graph: Leveraging Graph Neural Networks for Enhanced Explainability in NLP
von: Yáñez-Romero, Fabio, et al.
Veröffentlicht: (2025)
von: Yáñez-Romero, Fabio, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation
von: Bhattacharjee, Amrita, et al.
Veröffentlicht: (2024) -
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
von: Liu, Qihao, et al.
Veröffentlicht: (2025) -
Context-Enhanced Contrastive Search for Improved LLM Text Generation
von: Sen, Jaydip, et al.
Veröffentlicht: (2025) -
NLP Verification: Towards a General Methodology for Certifying Robustness
von: Casadio, Marco, et al.
Veröffentlicht: (2024) -
Muon is Scalable for LLM Training
von: Liu, Jingyuan, et al.
Veröffentlicht: (2025)