Navigating Semantic Relations: Challenges for Language Models in Abstract Common-Sense Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Gawin, Cole, Sun, Yidan, Kejriwal, Mayank |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Relational Schemata in BERT Are Inducible, Not Emergent: A Study of Performance vs. Competence in Language Models
por: Gawin, Cole
Publicado: (2025)
por: Gawin, Cole
Publicado: (2025)
Defining and Evaluating Decision and Composite Risk in Language Models Applied to Natural Language Inference
por: Shen, Ke, et al.
Publicado: (2024)
por: Shen, Ke, et al.
Publicado: (2024)
HALO: An Ontology for Representing and Categorizing Hallucinations in Large Language Models
por: Nananukul, Navapat, et al.
Publicado: (2023)
por: Nananukul, Navapat, et al.
Publicado: (2023)
Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models
por: Tang, Zhisheng, et al.
Publicado: (2024)
por: Tang, Zhisheng, et al.
Publicado: (2024)
GRASP: A Grid-Based Benchmark for Evaluating Commonsense Spatial Reasoning
por: Tang, Zhisheng, et al.
Publicado: (2024)
por: Tang, Zhisheng, et al.
Publicado: (2024)
Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models
por: Aravindan, Ashwath Vaithinathan, et al.
Publicado: (2026)
por: Aravindan, Ashwath Vaithinathan, et al.
Publicado: (2026)
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
por: Aravindan, Ashwath Vaithinathan, et al.
Publicado: (2026)
por: Aravindan, Ashwath Vaithinathan, et al.
Publicado: (2026)
SelECT-SQL: Self-correcting ensemble Chain-of-Thought for Text-to-SQL
por: Shen, Ke, et al.
Publicado: (2024)
por: Shen, Ke, et al.
Publicado: (2024)
An Evaluation of Estimative Uncertainty in Large Language Models
por: Tang, Zhisheng, et al.
Publicado: (2024)
por: Tang, Zhisheng, et al.
Publicado: (2024)
Is persona enough for personality? Using ChatGPT to reconstruct an agent's latent personality from simple descriptions
por: Ji, Yongyi, et al.
Publicado: (2024)
por: Ji, Yongyi, et al.
Publicado: (2024)
Exploring a Cognitive Architecture for Learning Arithmetic Equations
por: Gawin, Cole
Publicado: (2024)
por: Gawin, Cole
Publicado: (2024)
Theory Discovery in Social Networks: Automating ERGM Specification with Large Language Models
por: Sun, Yidan, et al.
Publicado: (2026)
por: Sun, Yidan, et al.
Publicado: (2026)
What is an "Abstract Reasoner"? Revisiting Experiments and Arguments about Large Language Models
por: Yun, Tian, et al.
Publicado: (2025)
por: Yun, Tian, et al.
Publicado: (2025)
Frame of Reference: Addressing the Challenges of Common Ground Representation in Situational Dialogs
por: Mohapatra, Biswesh, et al.
Publicado: (2026)
por: Mohapatra, Biswesh, et al.
Publicado: (2026)
Emergent Symbolic Mechanisms Support Abstract Reasoning in Large Language Models
por: Yang, Yukang, et al.
Publicado: (2025)
por: Yang, Yukang, et al.
Publicado: (2025)
Abstract Activation Spaces for Content-Invariant Reasoning in Large Language Models
por: Maraia, Gabriele, et al.
Publicado: (2026)
por: Maraia, Gabriele, et al.
Publicado: (2026)
IPPON: Common Sense Guided Informative Path Planning for Object Goal Navigation
por: Qu, Kaixian, et al.
Publicado: (2024)
por: Qu, Kaixian, et al.
Publicado: (2024)
An Analysis of Artificial Intelligence Adoption in NIH-Funded Research
por: Nananukul, Navapat, et al.
Publicado: (2026)
por: Nananukul, Navapat, et al.
Publicado: (2026)
Meaningful Learning: Enhancing Abstract Reasoning in Large Language Models via Generic Fact Guidance
por: Xiong, Kai, et al.
Publicado: (2024)
por: Xiong, Kai, et al.
Publicado: (2024)
From No to Know: Taxonomy, Challenges, and Opportunities for Negation Understanding in Multimodal Foundation Models
por: Vatsa, Mayank, et al.
Publicado: (2025)
por: Vatsa, Mayank, et al.
Publicado: (2025)
FRIDA to the Rescue! Analyzing Synthetic Data Effectiveness in Object-Based Common Sense Reasoning for Disaster Response
por: Shichman, Mollie, et al.
Publicado: (2025)
por: Shichman, Mollie, et al.
Publicado: (2025)
Cross-lingual Editing in Multilingual Language Models
por: Beniwal, Himanshu, et al.
Publicado: (2024)
por: Beniwal, Himanshu, et al.
Publicado: (2024)
Large Language Models as Common-Sense Heuristics
por: Borro, Andrey, et al.
Publicado: (2025)
por: Borro, Andrey, et al.
Publicado: (2025)
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models
por: Yeh, Cheng-Kai, et al.
Publicado: (2025)
por: Yeh, Cheng-Kai, et al.
Publicado: (2025)
CommonIT: Commonality-Aware Instruction Tuning for Large Language Models via Data Partitions
por: Rao, Jun, et al.
Publicado: (2024)
por: Rao, Jun, et al.
Publicado: (2024)
Do AI Models Perform Human-like Abstract Reasoning Across Modalities?
por: Beger, Claas, et al.
Publicado: (2025)
por: Beger, Claas, et al.
Publicado: (2025)
The Emergence of Abstract Thought in Large Language Models Beyond Any Language
por: Chen, Yuxin, et al.
Publicado: (2025)
por: Chen, Yuxin, et al.
Publicado: (2025)
Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges
por: Amjad, Husnain, et al.
Publicado: (2026)
por: Amjad, Husnain, et al.
Publicado: (2026)
Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning
por: Pinier, Christopher, et al.
Publicado: (2025)
por: Pinier, Christopher, et al.
Publicado: (2025)
Remember This Event That Year? Assessing Temporal Information and Reasoning in Large Language Models
por: Beniwal, Himanshu, et al.
Publicado: (2024)
por: Beniwal, Himanshu, et al.
Publicado: (2024)
LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
por: Zhou, Qianrui, et al.
Publicado: (2025)
por: Zhou, Qianrui, et al.
Publicado: (2025)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
por: Beniwal, Himanshu, et al.
Publicado: (2026)
por: Beniwal, Himanshu, et al.
Publicado: (2026)
Beyond the Star Rating: A Scalable Framework for Aspect-Based Sentiment Analysis Using LLMs and Text Classification
por: Patil, Vishal, et al.
Publicado: (2026)
por: Patil, Vishal, et al.
Publicado: (2026)
A Reliable Common-Sense Reasoning Socialbot Built Using LLMs and Goal-Directed ASP
por: Zeng, Yankai, et al.
Publicado: (2024)
por: Zeng, Yankai, et al.
Publicado: (2024)
Large Language Models as an Indirect Reasoner: Contrapositive and Contradiction for Automated Reasoning
por: Zhang, Yanfang, et al.
Publicado: (2024)
por: Zhang, Yanfang, et al.
Publicado: (2024)
Navigating the Concept Space of Language Models
por: Marcílio-Jr, Wilson E., et al.
Publicado: (2026)
por: Marcílio-Jr, Wilson E., et al.
Publicado: (2026)
No Universal Prompt: Unifying Reasoning through Adaptive Prompting for Temporal Table Reasoning
por: Rajgaria, Abhishek, et al.
Publicado: (2025)
por: Rajgaria, Abhishek, et al.
Publicado: (2025)
CausalARC: Abstract Reasoning with Causal World Models
por: Maasch, Jacqueline, et al.
Publicado: (2025)
por: Maasch, Jacqueline, et al.
Publicado: (2025)
Code-Driven Planning in Grid Worlds with Large Language Models
por: Aravindan, Ashwath Vaithinathan, et al.
Publicado: (2025)
por: Aravindan, Ashwath Vaithinathan, et al.
Publicado: (2025)
Generating Novelty in Open-World Multi-Agent Strategic Board Games
por: Kejriwal, Mayank, et al.
Publicado: (2025)
por: Kejriwal, Mayank, et al.
Publicado: (2025)
Ejemplares similares
-
Relational Schemata in BERT Are Inducible, Not Emergent: A Study of Performance vs. Competence in Language Models
por: Gawin, Cole
Publicado: (2025) -
Defining and Evaluating Decision and Composite Risk in Language Models Applied to Natural Language Inference
por: Shen, Ke, et al.
Publicado: (2024) -
HALO: An Ontology for Representing and Categorizing Hallucinations in Large Language Models
por: Nananukul, Navapat, et al.
Publicado: (2023) -
Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models
por: Tang, Zhisheng, et al.
Publicado: (2024) -
GRASP: A Grid-Based Benchmark for Evaluating Commonsense Spatial Reasoning
por: Tang, Zhisheng, et al.
Publicado: (2024)