A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia
Fuente:
arXiv
Guardado en:
| Autores principales: | Monea, Giovanni, Peyrard, Maxime, Josifoski, Martin, Chaudhary, Vishrav, Eisner, Jason, Kıcıman, Emre, Palangi, Hamid, Patra, Barun, West, Robert |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Agentic AI: The Era of Semantic Decoding
por: Peyrard, Maxime, et al.
Publicado: (2024)
por: Peyrard, Maxime, et al.
Publicado: (2024)
Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning
por: Geng, Saibo, et al.
Publicado: (2023)
por: Geng, Saibo, et al.
Publicado: (2023)
Interactive Evaluation of Large Language Models for Multi-Requirement Software Engineering Tasks
por: Rontogiannis, Dimitrios, et al.
Publicado: (2025)
por: Rontogiannis, Dimitrios, et al.
Publicado: (2025)
Evaluating Language Model Agency through Negotiations
por: Davidson, Tim R., et al.
Publicado: (2024)
por: Davidson, Tim R., et al.
Publicado: (2024)
Symbolic Autoencoding for Self-Supervised Sequence Learning
por: Amani, Mohammad Hossein, et al.
Publicado: (2024)
por: Amani, Mohammad Hossein, et al.
Publicado: (2024)
Do Llamas Work in English? On the Latent Language of Multilingual Transformers
por: Wendler, Chris, et al.
Publicado: (2024)
por: Wendler, Chris, et al.
Publicado: (2024)
A Practical Analysis of Human Alignment with *PO
por: Ahrabian, Kian, et al.
Publicado: (2024)
por: Ahrabian, Kian, et al.
Publicado: (2024)
Scaling Laws for Multilingual Language Models
por: He, Yifei, et al.
Publicado: (2024)
por: He, Yifei, et al.
Publicado: (2024)
Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers
por: Dumas, Clément, et al.
Publicado: (2024)
por: Dumas, Clément, et al.
Publicado: (2024)
S2-Attention: Hardware-Aware Context Sharding Among Attention Heads
por: Lin, Xihui, et al.
Publicado: (2024)
por: Lin, Xihui, et al.
Publicado: (2024)
Exploring Group and Symmetry Principles in Large Language Models
por: Imani, Shima, et al.
Publicado: (2024)
por: Imani, Shima, et al.
Publicado: (2024)
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
por: Kıcıman, Emre, et al.
Publicado: (2023)
por: Kıcıman, Emre, et al.
Publicado: (2023)
Sketch-Guided Constrained Decoding for Boosting Blackbox Large Language Models without Logit Access
por: Geng, Saibo, et al.
Publicado: (2024)
por: Geng, Saibo, et al.
Publicado: (2024)
Flows: Building Blocks of Reasoning and Collaborating AI
por: Josifoski, Martin, et al.
Publicado: (2023)
por: Josifoski, Martin, et al.
Publicado: (2023)
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
por: Matton, Katie, et al.
Publicado: (2025)
por: Matton, Katie, et al.
Publicado: (2025)
Meta-Statistical Learning: Supervised Learning of Statistical Estimators
por: Peyrard, Maxime, et al.
Publicado: (2025)
por: Peyrard, Maxime, et al.
Publicado: (2025)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
por: Méloux, Maxime, et al.
Publicado: (2025)
por: Méloux, Maxime, et al.
Publicado: (2025)
Controllable Context Sensitivity and the Knob Behind It
por: Minder, Julian, et al.
Publicado: (2024)
por: Minder, Julian, et al.
Publicado: (2024)
Modeling the Data-Generating Process is Necessary for Out-of-Distribution Generalization
por: Kaur, Jivat Neet, et al.
Publicado: (2022)
por: Kaur, Jivat Neet, et al.
Publicado: (2022)
The Dead Salmons of AI Interpretability
por: Méloux, Maxime, et al.
Publicado: (2025)
por: Méloux, Maxime, et al.
Publicado: (2025)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
por: Méloux, Maxime, et al.
Publicado: (2025)
por: Méloux, Maxime, et al.
Publicado: (2025)
Date Fragments: A Hidden Bottleneck of Tokenization for Temporal Reasoning
por: Bhatia, Gagan, et al.
Publicado: (2025)
por: Bhatia, Gagan, et al.
Publicado: (2025)
REFINER: Reasoning Feedback on Intermediate Representations
por: Paul, Debjit, et al.
Publicado: (2023)
por: Paul, Debjit, et al.
Publicado: (2023)
zip2zip: Inference-Time Adaptive Tokenization via Online Compression
por: Geng, Saibo, et al.
Publicado: (2025)
por: Geng, Saibo, et al.
Publicado: (2025)
TempGlitch: Evaluating Vision-Language Models for Temporal Glitch Detection in Gameplay Videos
por: Yu, Yakun, et al.
Publicado: (2026)
por: Yu, Yakun, et al.
Publicado: (2026)
GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language Models
por: Zhang, Zhibo, et al.
Publicado: (2024)
por: Zhang, Zhibo, et al.
Publicado: (2024)
Accelerating Language Model Workflows with Prompt Choreography
por: Bai, TJ, et al.
Publicado: (2025)
por: Bai, TJ, et al.
Publicado: (2025)
What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time?
por: Bhatia, Gagan, et al.
Publicado: (2026)
por: Bhatia, Gagan, et al.
Publicado: (2026)
LLMs Are In-Context Bandit Reinforcement Learners
por: Monea, Giovanni, et al.
Publicado: (2024)
por: Monea, Giovanni, et al.
Publicado: (2024)
EfficientXLang: Towards Improving Token Efficiency Through Cross-Lingual Reasoning
por: Ahuja, Sanchit, et al.
Publicado: (2025)
por: Ahuja, Sanchit, et al.
Publicado: (2025)
Open-Ended Video Game Glitch Detection with Agentic Reasoning and Temporal Grounding
por: Zheng, Muyang, et al.
Publicado: (2026)
por: Zheng, Muyang, et al.
Publicado: (2026)
Narcotweets: Social Media in Wartime
por: Monroy-Hernández, Andrés, et al.
Publicado: (2015)
por: Monroy-Hernández, Andrés, et al.
Publicado: (2015)
sPhinX: Sample Efficient Multilingual Instruction Fine-Tuning Through N-shot Guided Prompting
por: Ahuja, Sanchit, et al.
Publicado: (2024)
por: Ahuja, Sanchit, et al.
Publicado: (2024)
MIST: Mutual Information Estimation Via Supervised Training
por: Gritsai, German, et al.
Publicado: (2025)
por: Gritsai, German, et al.
Publicado: (2025)
Scaling Optimal LR Across Token Horizons
por: Bjorck, Johan, et al.
Publicado: (2024)
por: Bjorck, Johan, et al.
Publicado: (2024)
Point of View--Political Glitches
por: Shaw, Robert
Publicado: (2004)
por: Shaw, Robert
Publicado: (2004)
What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search
por: Zhang, Xinhao, et al.
Publicado: (2026)
por: Zhang, Xinhao, et al.
Publicado: (2026)
Domain Adaptation for Sustainable Soil Management using Causal and Contrastive Constraint Minimization
por: Sharma, Somya, et al.
Publicado: (2024)
por: Sharma, Somya, et al.
Publicado: (2024)
POROver: Improving Safety and Reducing Overrefusal in Large Language Models with Overgeneration and Preference Optimization
por: Karaman, Batuhan K., et al.
Publicado: (2024)
por: Karaman, Batuhan K., et al.
Publicado: (2024)
Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
por: Yi, Jingwei, et al.
Publicado: (2023)
por: Yi, Jingwei, et al.
Publicado: (2023)
Ejemplares similares
-
Agentic AI: The Era of Semantic Decoding
por: Peyrard, Maxime, et al.
Publicado: (2024) -
Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning
por: Geng, Saibo, et al.
Publicado: (2023) -
Interactive Evaluation of Large Language Models for Multi-Requirement Software Engineering Tasks
por: Rontogiannis, Dimitrios, et al.
Publicado: (2025) -
Evaluating Language Model Agency through Negotiations
por: Davidson, Tim R., et al.
Publicado: (2024) -
Symbolic Autoencoding for Self-Supervised Sequence Learning
por: Amani, Mohammad Hossein, et al.
Publicado: (2024)