Semantic Gravity Wells: Why Negative Constraints Backfire
Fuente:
arXiv
Salvato in:
| Autore principale: | Rana, Shailesh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Shape of Wisdom: Decision Trajectories in Language Models
di: Rana, Shailesh
Pubblicazione: (2026)
di: Rana, Shailesh
Pubblicazione: (2026)
When Chain-of-Thought Backfires: Evaluating Prompt Sensitivity in Medical Language Models
di: Sadanandan, Binesh, et al.
Pubblicazione: (2026)
di: Sadanandan, Binesh, et al.
Pubblicazione: (2026)
Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!
di: Zhou, Zhanhui, et al.
Pubblicazione: (2024)
di: Zhou, Zhanhui, et al.
Pubblicazione: (2024)
NCO: A Versatile Plug-in for Handling Negative Constraints in Decoding
di: Jin, Hyundong, et al.
Pubblicazione: (2026)
di: Jin, Hyundong, et al.
Pubblicazione: (2026)
Negation Triplet Extraction with Syntactic Dependency and Semantic Consistency
di: Shi, Yuchen, et al.
Pubblicazione: (2024)
di: Shi, Yuchen, et al.
Pubblicazione: (2024)
Enhancing Semantics in Multimodal Chain of Thought via Soft Negative Sampling
di: Zheng, Guangmin, et al.
Pubblicazione: (2024)
di: Zheng, Guangmin, et al.
Pubblicazione: (2024)
Semantic Adapter for Universal Text Embeddings: Diagnosing and Mitigating Negation Blindness to Enhance Universality
di: Cao, Hongliu
Pubblicazione: (2025)
di: Cao, Hongliu
Pubblicazione: (2025)
Mitigating Semantic Leakage in Cross-lingual Embeddings via Orthogonality Constraint
di: Ki, Dayeon, et al.
Pubblicazione: (2024)
di: Ki, Dayeon, et al.
Pubblicazione: (2024)
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
di: Duan, Shitong, et al.
Pubblicazione: (2024)
di: Duan, Shitong, et al.
Pubblicazione: (2024)
When Incentives Backfire, Data Stops Being Human
di: Santy, Sebastin, et al.
Pubblicazione: (2025)
di: Santy, Sebastin, et al.
Pubblicazione: (2025)
Alignment Backfire: Language-Dependent Reversal of Safety Interventions Across 16 Languages in LLM Multi-Agent Systems
di: Fukui, Hiroki
Pubblicazione: (2026)
di: Fukui, Hiroki
Pubblicazione: (2026)
Which bird does not have wings: Negative-constrained KGQA with Schema-guided Semantic Matching and Self-directed Refinement
di: Shim, Midan, et al.
Pubblicazione: (2026)
di: Shim, Midan, et al.
Pubblicazione: (2026)
Under Pressure: Emotional Framing Induces Measurable Behavioral Shifts and Structured Internal Geometry in Small Language Models
di: Usman, Rana Muhammad
Pubblicazione: (2026)
di: Usman, Rana Muhammad
Pubblicazione: (2026)
Semantic Integrity Constraints: Declarative Guardrails for AI-Augmented Data Processing Systems
di: Lee, Alexander W., et al.
Pubblicazione: (2025)
di: Lee, Alexander W., et al.
Pubblicazione: (2025)
Correcting Negative Bias in Large Language Models through Negative Attention Score Alignment
di: Yu, Sangwon, et al.
Pubblicazione: (2024)
di: Yu, Sangwon, et al.
Pubblicazione: (2024)
Semantic Anchors in In-Context Learning: Why Small LLMs Cannot Flip Their Labels
di: Kumar, Anantha Padmanaban Krishna
Pubblicazione: (2025)
di: Kumar, Anantha Padmanaban Krishna
Pubblicazione: (2025)
A Pseudo-Semantic Loss for Autoregressive Models with Logical Constraints
di: Ahmed, Kareem, et al.
Pubblicazione: (2023)
di: Ahmed, Kareem, et al.
Pubblicazione: (2023)
WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2024)
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2024)
Why are LLMs' abilities emergent?
di: Havlík, Vladimír
Pubblicazione: (2025)
di: Havlík, Vladimír
Pubblicazione: (2025)
Why Attend to Everything? Focus is the Key
di: Yao, Hengshuai, et al.
Pubblicazione: (2026)
di: Yao, Hengshuai, et al.
Pubblicazione: (2026)
Generating Diverse Negations from Affirmative Sentences
di: Vasquez, Darian Rodriguez, et al.
Pubblicazione: (2024)
di: Vasquez, Darian Rodriguez, et al.
Pubblicazione: (2024)
AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data
di: Alshaikh, Rana, et al.
Pubblicazione: (2025)
di: Alshaikh, Rana, et al.
Pubblicazione: (2025)
Fake Alignment: Are LLMs Really Aligned Well?
di: Wang, Yixu, et al.
Pubblicazione: (2023)
di: Wang, Yixu, et al.
Pubblicazione: (2023)
Reasoning Models Reason Well, Until They Don't
di: Rameshkumar, Revanth, et al.
Pubblicazione: (2025)
di: Rameshkumar, Revanth, et al.
Pubblicazione: (2025)
Why Slop Matters
di: Kommers, Cody, et al.
Pubblicazione: (2025)
di: Kommers, Cody, et al.
Pubblicazione: (2025)
Large-Scale Constraint Generation -- Can LLMs Parse Hundreds of Constraints?
di: Boffa, Matteo, et al.
Pubblicazione: (2025)
di: Boffa, Matteo, et al.
Pubblicazione: (2025)
The Impact of Negated Text on Hallucination with Large Language Models
di: Seo, Jaehyung, et al.
Pubblicazione: (2025)
di: Seo, Jaehyung, et al.
Pubblicazione: (2025)
How Well Do LLMs Understand Tunisian Arabic?
di: Mahdi, Mohamed
Pubblicazione: (2025)
di: Mahdi, Mohamed
Pubblicazione: (2025)
Why Chain of Thought Fails in Clinical Text Understanding
di: Wu, Jiageng, et al.
Pubblicazione: (2025)
di: Wu, Jiageng, et al.
Pubblicazione: (2025)
Why is constrained neural language generation particularly challenging?
di: Garbacea, Cristina, et al.
Pubblicazione: (2022)
di: Garbacea, Cristina, et al.
Pubblicazione: (2022)
Why Braking? Scenario Extraction and Reasoning Utilizing LLM
di: Wu, Yin, et al.
Pubblicazione: (2025)
di: Wu, Yin, et al.
Pubblicazione: (2025)
Carrot and Stick: Inducing Self-Motivation with Positive & Negative Feedback
di: Sohn, Jimin, et al.
Pubblicazione: (2024)
di: Sohn, Jimin, et al.
Pubblicazione: (2024)
Mitigating the Negative Impact of Over-association for Conversational Query Production
di: Wang, Ante, et al.
Pubblicazione: (2024)
di: Wang, Ante, et al.
Pubblicazione: (2024)
Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection
di: Pandit, Shrey, et al.
Pubblicazione: (2025)
di: Pandit, Shrey, et al.
Pubblicazione: (2025)
MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following
di: Lee, Jaeyun, et al.
Pubblicazione: (2026)
di: Lee, Jaeyun, et al.
Pubblicazione: (2026)
How Well Do Large Language Models Truly Ground?
di: Lee, Hyunji, et al.
Pubblicazione: (2023)
di: Lee, Hyunji, et al.
Pubblicazione: (2023)
Why Retrieval-Augmented Generation Fails: A Graph Perspective
di: Guo, Kai, et al.
Pubblicazione: (2026)
di: Guo, Kai, et al.
Pubblicazione: (2026)
Surgical Feature-Space Decomposition of LLMs: Why, When and How?
di: Chavan, Arnav, et al.
Pubblicazione: (2024)
di: Chavan, Arnav, et al.
Pubblicazione: (2024)
Look Within, Why LLMs Hallucinate: A Causal Perspective
di: Li, He, et al.
Pubblicazione: (2024)
di: Li, He, et al.
Pubblicazione: (2024)
Towards Minimal Targeted Updates of Language Models with Targeted Negative Training
di: Zhang, Lily H., et al.
Pubblicazione: (2024)
di: Zhang, Lily H., et al.
Pubblicazione: (2024)
Documenti analoghi
-
The Shape of Wisdom: Decision Trajectories in Language Models
di: Rana, Shailesh
Pubblicazione: (2026) -
When Chain-of-Thought Backfires: Evaluating Prompt Sensitivity in Medical Language Models
di: Sadanandan, Binesh, et al.
Pubblicazione: (2026) -
Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!
di: Zhou, Zhanhui, et al.
Pubblicazione: (2024) -
NCO: A Versatile Plug-in for Handling Negative Constraints in Decoding
di: Jin, Hyundong, et al.
Pubblicazione: (2026) -
Negation Triplet Extraction with Syntactic Dependency and Semantic Consistency
di: Shi, Yuchen, et al.
Pubblicazione: (2024)