Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Lang, Wan, Kaiyang, Liu, Wei, Wang, Chenxi, Song, Zirui, Xu, Zixiang, Wang, Yanbo, Stoyanov, Veselin, Chen, Xiuying |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Under the Shadow of Babel: How Language Shapes Reasoning in LLMs
by: Wang, Chenxi, et al.
Published: (2025)
by: Wang, Chenxi, et al.
Published: (2025)
Word Form Matters: LLMs' Semantic Reconstruction under Typoglycemia
by: Wang, Chenxi, et al.
Published: (2025)
by: Wang, Chenxi, et al.
Published: (2025)
Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies
by: Song, Zirui, et al.
Published: (2025)
by: Song, Zirui, et al.
Published: (2025)
Do LLMs "Feel"? Emotion Circuits Discovery and Control
by: Wang, Chenxi, et al.
Published: (2025)
by: Wang, Chenxi, et al.
Published: (2025)
DyFlow: Dynamic Workflow Framework for Agentic Reasoning
by: Wang, Yanbo, et al.
Published: (2025)
by: Wang, Yanbo, et al.
Published: (2025)
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
by: Song, Zirui, et al.
Published: (2025)
by: Song, Zirui, et al.
Published: (2025)
Evaluating and Mitigating Bias in AI-Based Medical Text Generation
by: Chen, Xiuying, et al.
Published: (2025)
by: Chen, Xiuying, et al.
Published: (2025)
The Cylindrical Representation Hypothesis for Language Model Steering
by: Gao, Lang, et al.
Published: (2026)
by: Gao, Lang, et al.
Published: (2026)
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
by: Xu, Zixiang, et al.
Published: (2025)
by: Xu, Zixiang, et al.
Published: (2025)
A Fano-Style Accuracy Upper Bound for LLM Single-Pass Reasoning in Multi-Hop QA
by: Wan, Kaiyang, et al.
Published: (2025)
by: Wan, Kaiyang, et al.
Published: (2025)
Flipping Knowledge Distillation: Leveraging Small Models' Expertise to Enhance LLMs in Text Matching
by: Li, Mingzhe, et al.
Published: (2025)
by: Li, Mingzhe, et al.
Published: (2025)
A Cognitive Writing Perspective for Constrained Long-Form Text Generation
by: Wan, Kaiyang, et al.
Published: (2025)
by: Wan, Kaiyang, et al.
Published: (2025)
When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection
by: Gao, Lang, et al.
Published: (2025)
by: Gao, Lang, et al.
Published: (2025)
Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models
by: Xu, Zixiang, et al.
Published: (2025)
by: Xu, Zixiang, et al.
Published: (2025)
From Individuals to Crowds: Dual-Level Public Response Prediction in Social Media
by: Zhang, Jinghui, et al.
Published: (2025)
by: Zhang, Jinghui, et al.
Published: (2025)
Adaptive Distraction: Probing LLM Contextual Robustness with Automated Tree Search
by: Wang, Yanbo, et al.
Published: (2025)
by: Wang, Yanbo, et al.
Published: (2025)
Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression
by: Zeng, Lingjie, et al.
Published: (2026)
by: Zeng, Lingjie, et al.
Published: (2026)
ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models
by: Song, Zirui, et al.
Published: (2025)
by: Song, Zirui, et al.
Published: (2025)
FineState-Bench: Benchmarking State-Conditioned Grounding for Fine-grained GUI State Setting
by: Ji, Fengxian, et al.
Published: (2026)
by: Ji, Fengxian, et al.
Published: (2026)
ServImage: An Image Generation and Editing Benchmark from Real-world Commercial Imaging Services
by: Ji, Fengxian, et al.
Published: (2026)
by: Ji, Fengxian, et al.
Published: (2026)
Adaptive KV-Cache Compression without Manually Setting Budget
by: Tang, Chenxia, et al.
Published: (2025)
by: Tang, Chenxia, et al.
Published: (2025)
Decoding Echo Chambers: LLM-Powered Simulations Revealing Polarization in Social Networks
by: Wang, Chenxi, et al.
Published: (2024)
by: Wang, Chenxi, et al.
Published: (2024)
Beyond Profile: From Surface-Level Facts to Deep Persona Simulation in LLMs
by: Wang, Zixiao, et al.
Published: (2025)
by: Wang, Zixiao, et al.
Published: (2025)
AI Security Beyond Core Domains: Resume Screening as a Case Study of Adversarial Vulnerabilities in Specialized LLM Applications
by: Mu, Honglin, et al.
Published: (2025)
by: Mu, Honglin, et al.
Published: (2025)
From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test
by: Dai, Xunlian, et al.
Published: (2025)
by: Dai, Xunlian, et al.
Published: (2025)
On the Effectiveness of LLMs for Manual Test Verifications
by: Peixoto, Myron David Lucena Campos, et al.
Published: (2024)
by: Peixoto, Myron David Lucena Campos, et al.
Published: (2024)
The aspect of brands and marketing development in pharmaceutical industry
by: Veselin Dickov
Published: (2012)
by: Veselin Dickov
Published: (2012)
Evaluating LLMs on Sequential API Call Through Automated Test Generation
by: Huang, Yuheng, et al.
Published: (2025)
by: Huang, Yuheng, et al.
Published: (2025)
Unsupervised Concept Vector Extraction for Bias Control in LLMs
by: Cyberey, Hannah, et al.
Published: (2025)
by: Cyberey, Hannah, et al.
Published: (2025)
Disentangle VAE for Molecular Generation
by: Wang, Yanbo, et al.
Published: (2022)
by: Wang, Yanbo, et al.
Published: (2022)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
On the Emergence of Weak-to-Strong Generalization: A Bias-Variance Perspective
by: Xu, Gengze, et al.
Published: (2025)
by: Xu, Gengze, et al.
Published: (2025)
Approximate Subgraph Matching with Neural Graph Representations and Reinforcement Learning
by: Li, Kaiyang, et al.
Published: (2026)
by: Li, Kaiyang, et al.
Published: (2026)
BlockSets: A Structured Visualization for Sets with Large Elements
by: Novakova, Neda, et al.
Published: (2025)
by: Novakova, Neda, et al.
Published: (2025)
The Stepwise Deception: Simulating the Evolution from True News to Fake News with LLM Agents
by: Liu, Yuhan, et al.
Published: (2024)
by: Liu, Yuhan, et al.
Published: (2024)
No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation
by: Yuan, Zhiqiang, et al.
Published: (2023)
by: Yuan, Zhiqiang, et al.
Published: (2023)
Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
by: Gao, Lang, et al.
Published: (2024)
by: Gao, Lang, et al.
Published: (2024)
MIST: Towards Multi-dimensional Implicit BiaS Evaluation of LLMs for Theory of Mind
by: Li, Yanlin, et al.
Published: (2025)
by: Li, Yanlin, et al.
Published: (2025)
Global Document Delivery, User Studies, and Service Evaluation: The Gateway Experience
by: Miller, Rush, et al.
Published: (2008)
by: Miller, Rush, et al.
Published: (2008)
Building on Efficient Foundations: Effectively Training LLMs with Structured Feedforward Layers
by: Wei, Xiuying, et al.
Published: (2024)
by: Wei, Xiuying, et al.
Published: (2024)
Similar Items
-
Under the Shadow of Babel: How Language Shapes Reasoning in LLMs
by: Wang, Chenxi, et al.
Published: (2025) -
Word Form Matters: LLMs' Semantic Reconstruction under Typoglycemia
by: Wang, Chenxi, et al.
Published: (2025) -
Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies
by: Song, Zirui, et al.
Published: (2025) -
Do LLMs "Feel"? Emotion Circuits Discovery and Control
by: Wang, Chenxi, et al.
Published: (2025) -
DyFlow: Dynamic Workflow Framework for Agentic Reasoning
by: Wang, Yanbo, et al.
Published: (2025)