The Geometry of Harmfulness in LLMs through Subconcept Probing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shah, McNair, Angeline, Saleena, Kumar, Adhitya Rajendra, Chheda, Naitik, Zhu, Kevin, Sharma, Vasu, O'Brien, Sean, Cai, Will |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
von: Chaturvedi, Isha, et al.
Veröffentlicht: (2025)
von: Chaturvedi, Isha, et al.
Veröffentlicht: (2025)
NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts
von: Gupta, Abhay, et al.
Veröffentlicht: (2025)
von: Gupta, Abhay, et al.
Veröffentlicht: (2025)
From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs
von: Yu, Stanley, et al.
Veröffentlicht: (2025)
von: Yu, Stanley, et al.
Veröffentlicht: (2025)
Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning
von: Lu, Leo, et al.
Veröffentlicht: (2025)
von: Lu, Leo, et al.
Veröffentlicht: (2025)
Beyond the Makerspace
von: Shivers-McNair, Ann
Veröffentlicht: (2021)
von: Shivers-McNair, Ann
Veröffentlicht: (2021)
A class of patch-use strategies. / James N. McNair
von: McNair, James N
Veröffentlicht: (1983)
von: McNair, James N
Veröffentlicht: (1983)
An Unsung Fashion Designer, Congo Square, Pinktoe Tarantulas, and Dead People: Expecting the Unexpected in Children's Literature
von: McNair, Jonda C.
Veröffentlicht: (2017)
von: McNair, Jonda C.
Veröffentlicht: (2017)
#WeNeedMirrorsAndWindows: Diverse Classroom Libraries for K-6 Students
von: McNair, Jonda C.
Veröffentlicht: (2016)
von: McNair, Jonda C.
Veröffentlicht: (2016)
Children as Social Butterflies: Navigating Belonging in a Diverse Swiss KindergartenBy UrsinaJaeger, New Brunswick, Camden, and Newark, New Jersey, London and Oxford: Rutgers University Press, 2025. ISBN: 978‐1‐9788‐3698‐3
von: Lynn J. McNair
Veröffentlicht: (2025)
von: Lynn J. McNair
Veröffentlicht: (2025)
Wayfinding through the AI wilderness: Mapping rhetorics of ChatGPT prompt writing on X (formerly Twitter) to promote critical AI literacies
von: Gupta, Anuj, et al.
Veröffentlicht: (2025)
von: Gupta, Anuj, et al.
Veröffentlicht: (2025)
SMAGDi: Socratic Multi Agent Interaction Graph Distillation for Efficient High Accuracy Reasoning
von: Aluru, Aayush, et al.
Veröffentlicht: (2025)
von: Aluru, Aayush, et al.
Veröffentlicht: (2025)
Sarc7: Evaluating Sarcasm Detection and Generation with Seven Types and Emotion-Informed Techniques
von: Xiong, Lang, et al.
Veröffentlicht: (2025)
von: Xiong, Lang, et al.
Veröffentlicht: (2025)
Pause-Tuning for Long-Context Comprehension: A Lightweight Approach to LLM Attention Recalibration
von: Begin, James, et al.
Veröffentlicht: (2025)
von: Begin, James, et al.
Veröffentlicht: (2025)
Pruning for Performance: Efficient Idiom and Metaphor Classification in Low-Resource Konkani Using mBERT
von: Do, Timothy, et al.
Veröffentlicht: (2025)
von: Do, Timothy, et al.
Veröffentlicht: (2025)
Interpreting the Latent Structure of Operator Precedence in Language Models
von: Yugeswardeenoo, Dharunish, et al.
Veröffentlicht: (2025)
von: Yugeswardeenoo, Dharunish, et al.
Veröffentlicht: (2025)
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
von: Chou, Cheng-Ting, et al.
Veröffentlicht: (2025)
von: Chou, Cheng-Ting, et al.
Veröffentlicht: (2025)
ERGO: Entropy-guided Resetting for Generation Optimization in Multi-turn Language Models
von: Khalid, Haziq Mohammad, et al.
Veröffentlicht: (2025)
von: Khalid, Haziq Mohammad, et al.
Veröffentlicht: (2025)
Universal Neurons in GPT-2: Emergence, Persistence, and Functional Impact
von: Nandan, Advey, et al.
Veröffentlicht: (2025)
von: Nandan, Advey, et al.
Veröffentlicht: (2025)
Adaptive Originality Filtering: Rejection Based Prompting and RiddleScore for Culturally Grounded Multilingual Riddle Generation
von: Le, Duy, et al.
Veröffentlicht: (2025)
von: Le, Duy, et al.
Veröffentlicht: (2025)
"But This Story of Mine Is Not Unique": A Review of Research on African American Children's Literature
von: Brooks, Wanda, et al.
Veröffentlicht: (2009)
von: Brooks, Wanda, et al.
Veröffentlicht: (2009)
Probing Audio-Generation Capabilities of Text-Based Language Models
von: Anbazhagan, Arjun Prasaath, et al.
Veröffentlicht: (2025)
von: Anbazhagan, Arjun Prasaath, et al.
Veröffentlicht: (2025)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
von: Csizmadia, Daniel, et al.
Veröffentlicht: (2025)
von: Csizmadia, Daniel, et al.
Veröffentlicht: (2025)
WOLF: Werewolf-based Observations for LLM Deception and Falsehoods
von: Agarwal, Mrinal, et al.
Veröffentlicht: (2025)
von: Agarwal, Mrinal, et al.
Veröffentlicht: (2025)
TRUTH DECAY: Quantifying Multi-Turn Sycophancy in Language Models
von: Liu, Joshua, et al.
Veröffentlicht: (2025)
von: Liu, Joshua, et al.
Veröffentlicht: (2025)
Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks
von: Yugeswardeenoo, Dharunish, et al.
Veröffentlicht: (2024)
von: Yugeswardeenoo, Dharunish, et al.
Veröffentlicht: (2024)
Deconstructing Bias: A Multifaceted Framework for Diagnosing Cultural and Compositional Inequities in Text-to-Image Generative Models
von: Said, Muna Numan, et al.
Veröffentlicht: (2025)
von: Said, Muna Numan, et al.
Veröffentlicht: (2025)
A Survey of Software-Defined Smart Grid Networks: Security Threats and Defense Techniques
von: Agnew, Dennis, et al.
Veröffentlicht: (2023)
von: Agnew, Dennis, et al.
Veröffentlicht: (2023)
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
von: Chintapatla, Ishant, et al.
Veröffentlicht: (2025)
von: Chintapatla, Ishant, et al.
Veröffentlicht: (2025)
Awareness and Usage of Digital Payment Gateway Among College Students in Kerala
von: Saleena. A P
Veröffentlicht: (2026)
von: Saleena. A P
Veröffentlicht: (2026)
Correcting Performance Estimation Bias in Imbalanced Classification with Minority Subconcepts
von: Maxson, Taylor, et al.
Veröffentlicht: (2026)
von: Maxson, Taylor, et al.
Veröffentlicht: (2026)
Rosetta-PL: Propositional Logic as a Benchmark for Large Language Model Reasoning
von: Baek, Shaun, et al.
Veröffentlicht: (2025)
von: Baek, Shaun, et al.
Veröffentlicht: (2025)
Advancing Uto-Aztecan Language Technologies: A Case Study on the Endangered Comanche Language
von: C, Jesus Alvarez, et al.
Veröffentlicht: (2025)
von: C, Jesus Alvarez, et al.
Veröffentlicht: (2025)
Probing localization properties of many-body Hamiltonians via an imaginary vector potential
von: O'Brien, Liam, et al.
Veröffentlicht: (2023)
von: O'Brien, Liam, et al.
Veröffentlicht: (2023)
A novel k-means clustering approach using two distance measures for Gaussian data
von: Gada, Naitik
Veröffentlicht: (2025)
von: Gada, Naitik
Veröffentlicht: (2025)
Pedagogical Reform at Primary Schools in Nepal: Examining the Child Centred Teaching
von: Shah, Rajendra Kumar
Veröffentlicht: (2020)
von: Shah, Rajendra Kumar
Veröffentlicht: (2020)
Error Reflection Prompting: Can Large Language Models Successfully Understand Errors?
von: Li, Jason, et al.
Veröffentlicht: (2025)
von: Li, Jason, et al.
Veröffentlicht: (2025)
CLEAR: Contrasting Textual Feedback with Experts and Amateurs for Reasoning
von: Rufail, Andrew, et al.
Veröffentlicht: (2025)
von: Rufail, Andrew, et al.
Veröffentlicht: (2025)
On the eve of the millennium : the future of democracy through an age of unreason / Conor Cruise O'Brien
von: O'Brien, Conor Cruise
von: O'Brien, Conor Cruise
Reproductive data for the western mosquitofish (Gambusia affinis)
von: Senior, Alistair McNair, et al.
Veröffentlicht: (2016)
von: Senior, Alistair McNair, et al.
Veröffentlicht: (2016)
Mass and length of female Gambusia affinis, and mass of her propagules and their genotypic sex
von: Senior, Alistair McNair, et al.
Veröffentlicht: (2016)
von: Senior, Alistair McNair, et al.
Veröffentlicht: (2016)
Ähnliche Einträge
-
Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
von: Chaturvedi, Isha, et al.
Veröffentlicht: (2025) -
NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts
von: Gupta, Abhay, et al.
Veröffentlicht: (2025) -
From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs
von: Yu, Stanley, et al.
Veröffentlicht: (2025) -
Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning
von: Lu, Leo, et al.
Veröffentlicht: (2025) -
Beyond the Makerspace
von: Shivers-McNair, Ann
Veröffentlicht: (2021)