Guardado en:
| Autores principales: | Luo, Yifan, Xu, Kangping, Lu, Yanzhen, Yuan, Yang, Yao, Andrew Chi-Chih |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2603.25187 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Automated Formal Verification of Backend Systems with LLMs
por: Xu, Kangping, et al.
Publicado: (2025)
por: Xu, Kangping, et al.
Publicado: (2025)
Group Representational Position Encoding
por: Zhang, Yifan, et al.
Publicado: (2025)
por: Zhang, Yifan, et al.
Publicado: (2025)
Meta Prompting for AI Systems
por: Zhang, Yifan, et al.
Publicado: (2023)
por: Zhang, Yifan, et al.
Publicado: (2023)
On the Diagram of Thought
por: Zhang, Yifan, et al.
Publicado: (2024)
por: Zhang, Yifan, et al.
Publicado: (2024)
Monadic Context Engineering
por: Zhang, Yifan, et al.
Publicado: (2025)
por: Zhang, Yifan, et al.
Publicado: (2025)
Augmenting Math Word Problems via Iterative Question Composing
por: Liu, Haoxiong, et al.
Publicado: (2024)
por: Liu, Haoxiong, et al.
Publicado: (2024)
On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
por: Zhang, Yifan, et al.
Publicado: (2025)
por: Zhang, Yifan, et al.
Publicado: (2025)
Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment
por: Zhang, Yifan, et al.
Publicado: (2024)
por: Zhang, Yifan, et al.
Publicado: (2024)
Tensor Product Attention Is All You Need
por: Zhang, Yifan, et al.
Publicado: (2025)
por: Zhang, Yifan, et al.
Publicado: (2025)
Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts
por: Zhang, Yifan, et al.
Publicado: (2024)
por: Zhang, Yifan, et al.
Publicado: (2024)
Activations as Features: Probing LLMs for Generalizable Essay Scoring Representations
por: Chi, Jinwei, et al.
Publicado: (2025)
por: Chi, Jinwei, et al.
Publicado: (2025)
Benchmarking Cognitive Domains for LLMs: Insights from Taiwanese Hakka Culture
por: Chang, Chen-Chi, et al.
Publicado: (2024)
por: Chang, Chen-Chi, et al.
Publicado: (2024)
Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions
por: Huang, Fan, et al.
Publicado: (2026)
por: Huang, Fan, et al.
Publicado: (2026)
The Self-Execution Benchmark: Measuring LLMs' Attempts to Overcome Their Lack of Self-Execution
por: Ezra, Elon, et al.
Publicado: (2025)
por: Ezra, Elon, et al.
Publicado: (2025)
From Human Cognition to Neural Activations: Probing the Computational Primitives of Spatial Reasoning in LLMs
por: An, Jiyuan, et al.
Publicado: (2026)
por: An, Jiyuan, et al.
Publicado: (2026)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
por: Agarwal, Dhruv, et al.
Publicado: (2025)
por: Agarwal, Dhruv, et al.
Publicado: (2025)
Mind the (Belief) Gap: Group Identity in the World of LLMs
por: Borah, Angana, et al.
Publicado: (2025)
por: Borah, Angana, et al.
Publicado: (2025)
Hallucination Detection with the Internal Layers of LLMs
por: Preiß, Martin
Publicado: (2025)
por: Preiß, Martin
Publicado: (2025)
Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State Uncertainty
por: Singh, Joykirat, et al.
Publicado: (2026)
por: Singh, Joykirat, et al.
Publicado: (2026)
MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining
por: Dong, Haoyu, et al.
Publicado: (2025)
por: Dong, Haoyu, et al.
Publicado: (2025)
Internal Knowledge Without External Expression: Probing the Generalization Boundary of a Classical Chinese Language Model
por: Chen, Jiuting, et al.
Publicado: (2026)
por: Chen, Jiuting, et al.
Publicado: (2026)
Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations
por: Wang, Yanli, et al.
Publicado: (2026)
por: Wang, Yanli, et al.
Publicado: (2026)
The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation
por: Xu, Rongwu, et al.
Publicado: (2023)
por: Xu, Rongwu, et al.
Publicado: (2023)
Evaluating Moral Beliefs across LLMs through a Pluralistic Framework
por: Liu, Xuelin, et al.
Publicado: (2024)
por: Liu, Xuelin, et al.
Publicado: (2024)
Jailbreak Instruction-Tuned LLMs via end-of-sentence MLP Re-weighting
por: Luo, Yifan, et al.
Publicado: (2024)
por: Luo, Yifan, et al.
Publicado: (2024)
Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals
por: Clymer, Joshua, et al.
Publicado: (2024)
por: Clymer, Joshua, et al.
Publicado: (2024)
CulturalTeaming: AI-Assisted Interactive Red-Teaming for Challenging LLMs' (Lack of) Multicultural Knowledge
por: Chiu, Yu Ying, et al.
Publicado: (2024)
por: Chiu, Yu Ying, et al.
Publicado: (2024)
ReProbe: Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
por: Ni, Jingwei, et al.
Publicado: (2025)
por: Ni, Jingwei, et al.
Publicado: (2025)
Finding the Cracks: Improving LLMs Reasoning with Paraphrastic Probing and Consistency Verification
por: Shi, Weili, et al.
Publicado: (2026)
por: Shi, Weili, et al.
Publicado: (2026)
Agentic Recommender System with Hierarchical Belief-State Memory
por: Shen, Xiang, et al.
Publicado: (2026)
por: Shen, Xiang, et al.
Publicado: (2026)
Using LLMs to Model the Beliefs and Preferences of Targeted Populations
por: Namikoshi, Keiichi, et al.
Publicado: (2024)
por: Namikoshi, Keiichi, et al.
Publicado: (2024)
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
por: Chen, Justin Chih-Yao, et al.
Publicado: (2023)
por: Chen, Justin Chih-Yao, et al.
Publicado: (2023)
Probing LLM Hallucination from Within: Perturbation-Driven Approach via Internal Knowledge
por: Lee, Seongmin, et al.
Publicado: (2024)
por: Lee, Seongmin, et al.
Publicado: (2024)
Fundamental Problems With Model Editing: How Should Rational Belief Revision Work in LLMs?
por: Hase, Peter, et al.
Publicado: (2024)
por: Hase, Peter, et al.
Publicado: (2024)
FairBelief -- Assessing Harmful Beliefs in Language Models
por: Setzu, Mattia, et al.
Publicado: (2024)
por: Setzu, Mattia, et al.
Publicado: (2024)
Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs
por: Kale, Sahil
Publicado: (2025)
por: Kale, Sahil
Publicado: (2025)
Beyond Words: A Latent Memory Approach to Internal Reasoning in LLMs
por: Orlicki, José I.
Publicado: (2025)
por: Orlicki, José I.
Publicado: (2025)
HeartBench: Probing Core Dimensions of Anthropomorphic Intelligence in LLMs
por: Liu, Jiaxin, et al.
Publicado: (2025)
por: Liu, Jiaxin, et al.
Publicado: (2025)
Impoverished Language Technology: The Lack of (Social) Class in NLP
por: Curry, Amanda Cercas, et al.
Publicado: (2024)
por: Curry, Amanda Cercas, et al.
Publicado: (2024)
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
por: Xu, Wanghan, et al.
Publicado: (2025)
por: Xu, Wanghan, et al.
Publicado: (2025)
Ejemplares similares
-
Towards Automated Formal Verification of Backend Systems with LLMs
por: Xu, Kangping, et al.
Publicado: (2025) -
Group Representational Position Encoding
por: Zhang, Yifan, et al.
Publicado: (2025) -
Meta Prompting for AI Systems
por: Zhang, Yifan, et al.
Publicado: (2023) -
On the Diagram of Thought
por: Zhang, Yifan, et al.
Publicado: (2024) -
Monadic Context Engineering
por: Zhang, Yifan, et al.
Publicado: (2025)