Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?
Fuente:
arXiv
Salvato in:
| Autori principali: | Shen, Siqi, Singh, Mehar, Logeswaran, Lajanugen, Lee, Moontae, Lee, Honglak, Mihalcea, Rada |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense
di: Shen, Siqi, et al.
Pubblicazione: (2024)
di: Shen, Siqi, et al.
Pubblicazione: (2024)
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
di: Khalifa, Muhammad, et al.
Pubblicazione: (2023)
di: Khalifa, Muhammad, et al.
Pubblicazione: (2023)
Small Language Models Need Strong Verifiers to Self-Correct Reasoning
di: Zhang, Yunxiang, et al.
Pubblicazione: (2024)
di: Zhang, Yunxiang, et al.
Pubblicazione: (2024)
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
di: Zhang, Lechen, et al.
Pubblicazione: (2025)
di: Zhang, Lechen, et al.
Pubblicazione: (2025)
Process Reward Models That Think
di: Khalifa, Muhammad, et al.
Pubblicazione: (2025)
di: Khalifa, Muhammad, et al.
Pubblicazione: (2025)
SPRIG: Improving Large Language Model Performance by System Prompt Optimization
di: Zhang, Lechen, et al.
Pubblicazione: (2024)
di: Zhang, Lechen, et al.
Pubblicazione: (2024)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
di: Zheng, Mingqian, et al.
Pubblicazione: (2023)
di: Zheng, Mingqian, et al.
Pubblicazione: (2023)
Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
di: Khalifa, Muhammad, et al.
Pubblicazione: (2026)
di: Khalifa, Muhammad, et al.
Pubblicazione: (2026)
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
di: Kim, Jaekyeom, et al.
Pubblicazione: (2024)
di: Kim, Jaekyeom, et al.
Pubblicazione: (2024)
Visual Test-time Scaling for GUI Agent Grounding
di: Luo, Tiange, et al.
Pubblicazione: (2025)
di: Luo, Tiange, et al.
Pubblicazione: (2025)
AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents
di: Fu, Yao, et al.
Pubblicazione: (2024)
di: Fu, Yao, et al.
Pubblicazione: (2024)
You don't need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments
di: Shu, Bangzhao, et al.
Pubblicazione: (2023)
di: Shu, Bangzhao, et al.
Pubblicazione: (2023)
Selective LoRA for Visual Tokens and Attention Heads
di: Luo, Tiange, et al.
Pubblicazione: (2025)
di: Luo, Tiange, et al.
Pubblicazione: (2025)
Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions
di: Borah, Angana, et al.
Pubblicazione: (2024)
di: Borah, Angana, et al.
Pubblicazione: (2024)
The Generation Gap: Exploring Age Bias in the Value Systems of Large Language Models
di: Liu, Siyang, et al.
Pubblicazione: (2024)
di: Liu, Siyang, et al.
Pubblicazione: (2024)
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
di: Jang, Yunseok, et al.
Pubblicazione: (2025)
di: Jang, Yunseok, et al.
Pubblicazione: (2025)
Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions
di: Deng, Naihao, et al.
Pubblicazione: (2025)
di: Deng, Naihao, et al.
Pubblicazione: (2025)
Persuasion at Play: Understanding Misinformation Dynamics in Demographic-Aware Human-LLM Interactions
di: Borah, Angana, et al.
Pubblicazione: (2025)
di: Borah, Angana, et al.
Pubblicazione: (2025)
Rethinking Table Instruction Tuning
di: Deng, Naihao, et al.
Pubblicazione: (2025)
di: Deng, Naihao, et al.
Pubblicazione: (2025)
Scaling Web Agent Training through Automatic Data Generation and Fine-grained Evaluation
di: Logeswaran, Lajanugen, et al.
Pubblicazione: (2026)
di: Logeswaran, Lajanugen, et al.
Pubblicazione: (2026)
Towards Algorithmic Fidelity: Mental Health Representation across Demographics in Synthetic vs. Human-generated Data
di: Mori, Shinka, et al.
Pubblicazione: (2024)
di: Mori, Shinka, et al.
Pubblicazione: (2024)
Benchmarking and Improving LLM Robustness for Personalized Generation
di: Okite, Chimaobi, et al.
Pubblicazione: (2025)
di: Okite, Chimaobi, et al.
Pubblicazione: (2025)
Cross-cultural Inspiration Detection and Analysis in Real and LLM-generated Social Media Data
di: Ignat, Oana, et al.
Pubblicazione: (2024)
di: Ignat, Oana, et al.
Pubblicazione: (2024)
The Curious Case of Curiosity across Human Cultures and LLMs
di: Borah, Angana, et al.
Pubblicazione: (2025)
di: Borah, Angana, et al.
Pubblicazione: (2025)
The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models
di: Arif, Samee, et al.
Pubblicazione: (2026)
di: Arif, Samee, et al.
Pubblicazione: (2026)
Towards Region-aware Bias Evaluation Metrics
di: Borah, Angana, et al.
Pubblicazione: (2024)
di: Borah, Angana, et al.
Pubblicazione: (2024)
One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety
di: Arif, Samee, et al.
Pubblicazione: (2026)
di: Arif, Samee, et al.
Pubblicazione: (2026)
Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models
di: Piedrahita, David Guzman, et al.
Pubblicazione: (2025)
di: Piedrahita, David Guzman, et al.
Pubblicazione: (2025)
Whose wife is it anyway? Assessing bias against same-gender relationships in machine translation
di: Stewart, Ian, et al.
Pubblicazione: (2024)
di: Stewart, Ian, et al.
Pubblicazione: (2024)
DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
di: Han, Janghoon, et al.
Pubblicazione: (2025)
di: Han, Janghoon, et al.
Pubblicazione: (2025)
Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal Models
di: Nwatu, Joan, et al.
Pubblicazione: (2024)
di: Nwatu, Joan, et al.
Pubblicazione: (2024)
MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
di: Zhang, Yunxiang, et al.
Pubblicazione: (2025)
di: Zhang, Yunxiang, et al.
Pubblicazione: (2025)
Mind the (Belief) Gap: Group Identity in the World of LLMs
di: Borah, Angana, et al.
Pubblicazione: (2025)
di: Borah, Angana, et al.
Pubblicazione: (2025)
MAiDE-up: Multilingual Deception Detection of GPT-generated Hotel Reviews
di: Ignat, Oana, et al.
Pubblicazione: (2024)
di: Ignat, Oana, et al.
Pubblicazione: (2024)
Patient-Centered RAG for Oncology Visit Aid Following the Ottawa Decision Guide
di: Liu, Siyang, et al.
Pubblicazione: (2025)
di: Liu, Siyang, et al.
Pubblicazione: (2025)
Towards Dog Bark Decoding: Leveraging Human Speech Processing for Automated Bark Classification
di: Abzaliev, Artem, et al.
Pubblicazione: (2024)
di: Abzaliev, Artem, et al.
Pubblicazione: (2024)
Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs
di: Cha, Sungmin, et al.
Pubblicazione: (2024)
di: Cha, Sungmin, et al.
Pubblicazione: (2024)
Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents
di: Piatti, Giorgio, et al.
Pubblicazione: (2024)
di: Piatti, Giorgio, et al.
Pubblicazione: (2024)
Value Drifts: Tracing Value Alignment During LLM Post-Training
di: Bhatia, Mehar, et al.
Pubblicazione: (2025)
di: Bhatia, Mehar, et al.
Pubblicazione: (2025)
Interactive and Expressive Code-Augmented Planning with Large Language Models
di: Liu, Anthony Z., et al.
Pubblicazione: (2024)
di: Liu, Anthony Z., et al.
Pubblicazione: (2024)
Documenti analoghi
-
Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense
di: Shen, Siqi, et al.
Pubblicazione: (2024) -
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
di: Khalifa, Muhammad, et al.
Pubblicazione: (2023) -
Small Language Models Need Strong Verifiers to Self-Correct Reasoning
di: Zhang, Yunxiang, et al.
Pubblicazione: (2024) -
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
di: Zhang, Lechen, et al.
Pubblicazione: (2025) -
Process Reward Models That Think
di: Khalifa, Muhammad, et al.
Pubblicazione: (2025)