TEMPER: Testing Emotional Perturbation in Quantitative Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Dokme, Atahan, Reichman, Benjamin, Heck, Larry |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Emotions Where Art Thou: Understanding and Characterizing the Emotional Latent Space of Large Language Models
por: Reichman, Benjamin, et al.
Publicado: (2025)
por: Reichman, Benjamin, et al.
Publicado: (2025)
Emotion is Not Just a Label: Latent Emotional Factors in LLM Processing
por: Reichman, Benjamin, et al.
Publicado: (2026)
por: Reichman, Benjamin, et al.
Publicado: (2026)
Reading with Intent -- Neutralizing Intent
por: Reichman, Benjamin, et al.
Publicado: (2025)
por: Reichman, Benjamin, et al.
Publicado: (2025)
Emotional RAG LLMs: Reading Comprehension for the Open Internet
por: Reichman, Benjamin, et al.
Publicado: (2024)
por: Reichman, Benjamin, et al.
Publicado: (2024)
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos
por: Reichman, Benjamin, et al.
Publicado: (2025)
por: Reichman, Benjamin, et al.
Publicado: (2025)
Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders
por: Dokme, Atahan, et al.
Publicado: (2026)
por: Dokme, Atahan, et al.
Publicado: (2026)
Weight-of-Thought Reasoning: Exploring Neural Network Weights for Enhanced LLM Reasoning
por: Punjwani, Saif, et al.
Publicado: (2025)
por: Punjwani, Saif, et al.
Publicado: (2025)
Dense Passage Retrieval: Is it Retrieving?
por: Reichman, Benjamin, et al.
Publicado: (2024)
por: Reichman, Benjamin, et al.
Publicado: (2024)
SensorQA: A Question Answering Benchmark for Daily-Life Monitoring
por: Reichman, Benjamin, et al.
Publicado: (2025)
por: Reichman, Benjamin, et al.
Publicado: (2025)
UNLEARN Efficient Removal of Knowledge in Large Language Models
por: Lizzo, Tyler, et al.
Publicado: (2024)
por: Lizzo, Tyler, et al.
Publicado: (2024)
A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
por: Karagoz, Atahan
Publicado: (2026)
por: Karagoz, Atahan
Publicado: (2026)
Allo-AVA: A Large-Scale Multimodal Conversational AI Dataset for Allocentric Avatar Gesture Animation
por: Punjwani, Saif, et al.
Publicado: (2024)
por: Punjwani, Saif, et al.
Publicado: (2024)
SensorChat: Answering Qualitative and Quantitative Questions during Long-Term Multimodal Sensor Interactions
por: Yu, Xiaofan, et al.
Publicado: (2025)
por: Yu, Xiaofan, et al.
Publicado: (2025)
Large Body Language Models
por: Punjwani, Saif, et al.
Publicado: (2024)
por: Punjwani, Saif, et al.
Publicado: (2024)
iTBLS: A Dataset of Interactive Conversations Over Tabular Information
por: Sundar, Anirudh, et al.
Publicado: (2024)
por: Sundar, Anirudh, et al.
Publicado: (2024)
Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm
por: Gupta, Akshat, et al.
Publicado: (2024)
por: Gupta, Akshat, et al.
Publicado: (2024)
Are Human Conversations Special? A Large Language Model Perspective
por: Jawale, Toshish, et al.
Publicado: (2024)
por: Jawale, Toshish, et al.
Publicado: (2024)
cPAPERS: A Dataset of Situated and Multimodal Interactive Conversations in Scientific Papers
por: Sundar, Anirudh, et al.
Publicado: (2024)
por: Sundar, Anirudh, et al.
Publicado: (2024)
Detecting Emotional Incongruity of Sarcasm by Commonsense Reasoning
por: Qiu, Ziqi, et al.
Publicado: (2024)
por: Qiu, Ziqi, et al.
Publicado: (2024)
Towards a Generative Approach for Emotion Detection and Reasoning
por: Bhaumik, Ankita, et al.
Publicado: (2024)
por: Bhaumik, Ankita, et al.
Publicado: (2024)
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
por: Liu, Xiao, et al.
Publicado: (2024)
por: Liu, Xiao, et al.
Publicado: (2024)
Understanding Artificial Theory of Mind: Perturbed Tasks and Reasoning in Large Language Models
por: Nickel, Christian, et al.
Publicado: (2026)
por: Nickel, Christian, et al.
Publicado: (2026)
Language Models can perform Single-Utterance Self-Correction of Perturbed Reasoning
por: Silver, Sam, et al.
Publicado: (2025)
por: Silver, Sam, et al.
Publicado: (2025)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
por: Huang, Kaixuan, et al.
Publicado: (2025)
por: Huang, Kaixuan, et al.
Publicado: (2025)
Norm Growth and Stability Challenges in Localized Sequential Knowledge Editing
por: Gupta, Akshat, et al.
Publicado: (2025)
por: Gupta, Akshat, et al.
Publicado: (2025)
How Do Answer Tokens Read Reasoning Traces? Self-Reading Patterns in Thinking LLMs for Quantitative Reasoning
por: Chen, Haoyang, et al.
Publicado: (2026)
por: Chen, Haoyang, et al.
Publicado: (2026)
EmoLLM: Appraisal-Grounded Cognitive-Emotional Co-Reasoning in Large Language Models
por: Zhang, Yifei, et al.
Publicado: (2026)
por: Zhang, Yifei, et al.
Publicado: (2026)
Entropy-Gated Branching for Efficient Test-Time Reasoning
por: Li, Xianzhi, et al.
Publicado: (2025)
por: Li, Xianzhi, et al.
Publicado: (2025)
Nonsense Helps: Prompt Space Perturbation Broadens Reasoning Exploration
por: Huang, Langlin, et al.
Publicado: (2026)
por: Huang, Langlin, et al.
Publicado: (2026)
Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning
por: Hu, Zhiyuan, et al.
Publicado: (2026)
por: Hu, Zhiyuan, et al.
Publicado: (2026)
LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling
por: Xiao, Yang, et al.
Publicado: (2025)
por: Xiao, Yang, et al.
Publicado: (2025)
Logical Reasoning with Outcome Reward Models for Test-Time Scaling
por: Thatikonda, Ramya Keerthy, et al.
Publicado: (2025)
por: Thatikonda, Ramya Keerthy, et al.
Publicado: (2025)
Precedent-Informed Reasoning: Mitigating Overthinking in Large Reasoning Models via Test-Time Precedent Learning
por: Wang, Qianyue, et al.
Publicado: (2026)
por: Wang, Qianyue, et al.
Publicado: (2026)
Characterizing the Robustness of Black-Box LLM Planners Under Perturbed Observations with Adaptive Stress Testing
por: Chakraborty, Neeloy, et al.
Publicado: (2025)
por: Chakraborty, Neeloy, et al.
Publicado: (2025)
Enhancing Depression Detection with Chain-of-Thought Prompting: From Emotion to Reasoning Using Large Language Models
por: Teng, Shiyu, et al.
Publicado: (2025)
por: Teng, Shiyu, et al.
Publicado: (2025)
TempPerturb-Eval: On the Joint Effects of Internal Temperature and External Perturbations in RAG Robustness
por: Zhou, Yongxin, et al.
Publicado: (2025)
por: Zhou, Yongxin, et al.
Publicado: (2025)
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
por: Zhou, Jin Peng, et al.
Publicado: (2024)
por: Zhou, Jin Peng, et al.
Publicado: (2024)
Hypothesis Testing Prompting Improves Deductive Reasoning in Large Language Models
por: Li, Yitian, et al.
Publicado: (2024)
por: Li, Yitian, et al.
Publicado: (2024)
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
por: Kuratov, Yuri, et al.
Publicado: (2024)
por: Kuratov, Yuri, et al.
Publicado: (2024)
MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
por: Chen, Hongwei, et al.
Publicado: (2025)
por: Chen, Hongwei, et al.
Publicado: (2025)
Ejemplares similares
-
Emotions Where Art Thou: Understanding and Characterizing the Emotional Latent Space of Large Language Models
por: Reichman, Benjamin, et al.
Publicado: (2025) -
Emotion is Not Just a Label: Latent Emotional Factors in LLM Processing
por: Reichman, Benjamin, et al.
Publicado: (2026) -
Reading with Intent -- Neutralizing Intent
por: Reichman, Benjamin, et al.
Publicado: (2025) -
Emotional RAG LLMs: Reading Comprehension for the Open Internet
por: Reichman, Benjamin, et al.
Publicado: (2024) -
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos
por: Reichman, Benjamin, et al.
Publicado: (2025)