Evaluating Contrast Localizer for Identifying Causal Units in Social & Mathematical Tasks in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Jamaa, Yassine, AlKhamissi, Badr, Ghosh, Satrajit, Schrimpf, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The LLM Language Network: A Neuroscientific Approach for Identifying Causally Task-Relevant Units
by: AlKhamissi, Badr, et al.
Published: (2024)
by: AlKhamissi, Badr, et al.
Published: (2024)
Rosetta Stone at KSAA-RD Shared Task: A Hop From Language Modeling To Word--Definition Alignment
by: ElBakry, Ahmed, et al.
Published: (2023)
by: ElBakry, Ahmed, et al.
Published: (2023)
Brain-Like Language Processing via a Shallow Untrained Multihead Attention Network
by: AlKhamissi, Badr, et al.
Published: (2024)
by: AlKhamissi, Badr, et al.
Published: (2024)
Inducing Dyslexia in Vision Language Models
by: Honarmand, Melika, et al.
Published: (2025)
by: Honarmand, Melika, et al.
Published: (2025)
Investigating Cultural Alignment of Large Language Models
by: AlKhamissi, Badr, et al.
Published: (2024)
by: AlKhamissi, Badr, et al.
Published: (2024)
A Context-Contrastive Inference Approach To Partial Diacritization
by: ElNokrashy, Muhammad, et al.
Published: (2024)
by: ElNokrashy, Muhammad, et al.
Published: (2024)
Hire Your Anthropologist! Rethinking Culture Benchmarks Through an Anthropological Lens
by: AlKhamissi, Mai, et al.
Published: (2025)
by: AlKhamissi, Mai, et al.
Published: (2025)
From Language to Cognition: How LLMs Outgrow the Human Language Network
by: AlKhamissi, Badr, et al.
Published: (2025)
by: AlKhamissi, Badr, et al.
Published: (2025)
Rational Metareasoning for Large Language Models
by: De Sabbata, C. Nicolò, et al.
Published: (2024)
by: De Sabbata, C. Nicolò, et al.
Published: (2024)
Instruction-tuning Aligns LLMs to the Human Brain
by: Aw, Khai Loong, et al.
Published: (2023)
by: Aw, Khai Loong, et al.
Published: (2023)
MIRAGE: Adaptive Multimodal Gating for Whole-Brain fMRI Encoding
by: Gokce, Abdulkadir, et al.
Published: (2026)
by: Gokce, Abdulkadir, et al.
Published: (2026)
TopoLM: brain-like spatio-functional organization in a topographic language model
by: Rathi, Neil, et al.
Published: (2024)
by: Rathi, Neil, et al.
Published: (2024)
Large Language Models Align with the Human Brain during Creative Thinking
by: Ismayilzada, Mete, et al.
Published: (2026)
by: Ismayilzada, Mete, et al.
Published: (2026)
Dreaming Out Loud: A Self-Synthesis Approach For Training Vision-Language Models With Developmentally Plausible Data
by: AlKhamissi, Badr, et al.
Published: (2024)
by: AlKhamissi, Badr, et al.
Published: (2024)
Depth-Wise Attention (DWAtt): A Layer Fusion Method for Data-Efficient Classification
by: ElNokrashy, Muhammad, et al.
Published: (2022)
by: ElNokrashy, Muhammad, et al.
Published: (2022)
"Flex Tape Can't Fix That": Bias and Misinformation in Edited Language Models
by: Halevy, Karina, et al.
Published: (2024)
by: Halevy, Karina, et al.
Published: (2024)
STRUCTSENSE: A Task-Agnostic Agentic Framework for Structured Information Extraction with Human-In-The-Loop Evaluation and Benchmarking
by: Chhetri, Tek Raj, et al.
Published: (2025)
by: Chhetri, Tek Raj, et al.
Published: (2025)
Mathify: Evaluating Large Language Models on Mathematical Problem Solving Tasks
by: Anand, Avinash, et al.
Published: (2024)
by: Anand, Avinash, et al.
Published: (2024)
Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models
by: Arad, Dana, et al.
Published: (2025)
by: Arad, Dana, et al.
Published: (2025)
Khattat: Enhancing Readability and Concept Representation of Semantic Typography
by: Hussein, Ahmed, et al.
Published: (2024)
by: Hussein, Ahmed, et al.
Published: (2024)
Role-Playing Evaluation for Large Language Models
by: Boudouri, Yassine El, et al.
Published: (2025)
by: Boudouri, Yassine El, et al.
Published: (2025)
Eliciting Causal Abilities in Large Language Models for Reasoning Tasks
by: Wang, Yajing, et al.
Published: (2024)
by: Wang, Yajing, et al.
Published: (2024)
Evaluating Large Language Models for Health-Related Text Classification Tasks with Public Social Media Data
by: Guo, Yuting, et al.
Published: (2024)
by: Guo, Yuting, et al.
Published: (2024)
Identifying and Mitigating Social Bias Knowledge in Language Models
by: Chen, Ruizhe, et al.
Published: (2024)
by: Chen, Ruizhe, et al.
Published: (2024)
The Qiyas Benchmark: Measuring ChatGPT Mathematical and Language Understanding in Arabic
by: Al-Khalifa, Shahad, et al.
Published: (2024)
by: Al-Khalifa, Shahad, et al.
Published: (2024)
Causal Evaluation of Language Models
by: Chen, Sirui, et al.
Published: (2024)
by: Chen, Sirui, et al.
Published: (2024)
Identifying Multiple Personalities in Large Language Models with External Evaluation
by: Song, Xiaoyang, et al.
Published: (2024)
by: Song, Xiaoyang, et al.
Published: (2024)
Evaluating the Efficacy of Large Language Models in Identifying Phishing Attempts
by: Patel, Het, et al.
Published: (2024)
by: Patel, Het, et al.
Published: (2024)
Introducing HALC: A general pipeline for finding optimal prompting strategies for automated coding with LLMs in the computational social sciences
by: Reich, Andreas, et al.
Published: (2025)
by: Reich, Andreas, et al.
Published: (2025)
Underspecification in Language Modeling Tasks: A Causality-Informed Study of Gendered Pronoun Resolution
by: McMilin, Emily
Published: (2022)
by: McMilin, Emily
Published: (2022)
Evaluating Ill-Defined Tasks in Large Language Models
by: Zhou, Yi, et al.
Published: (2026)
by: Zhou, Yi, et al.
Published: (2026)
Evaluating Large Language Models for Abstract Evaluation Tasks: An Empirical Study
by: Liu, Yinuo, et al.
Published: (2026)
by: Liu, Yinuo, et al.
Published: (2026)
Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges
by: Amjad, Husnain, et al.
Published: (2026)
by: Amjad, Husnain, et al.
Published: (2026)
Large Language Models to Identify Social Determinants of Health in Electronic Health Records
by: Guevara, Marco, et al.
Published: (2023)
by: Guevara, Marco, et al.
Published: (2023)
Reward Models Identify Consistency, Not Causality
by: Xu, Yuhui, et al.
Published: (2025)
by: Xu, Yuhui, et al.
Published: (2025)
LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
by: Saad-Falcon, Jon, et al.
Published: (2024)
by: Saad-Falcon, Jon, et al.
Published: (2024)
A Comparative Study of Task Adaptation Techniques of Large Language Models for Identifying Sustainable Development Goals
by: Cadeddu, Andrea, et al.
Published: (2025)
by: Cadeddu, Andrea, et al.
Published: (2025)
Evaluating Large Language Models for Real-World Engineering Tasks
by: Heesch, Rene, et al.
Published: (2025)
by: Heesch, Rene, et al.
Published: (2025)
M3Kang: Evaluating Multilingual Multimodal Mathematical Reasoning in Vision-Language Models
by: Torres-Camps, Aleix, et al.
Published: (2026)
by: Torres-Camps, Aleix, et al.
Published: (2026)
Large Language Model enabled Mathematical Modeling
by: Zhang, Guoyun
Published: (2025)
by: Zhang, Guoyun
Published: (2025)
Similar Items
-
The LLM Language Network: A Neuroscientific Approach for Identifying Causally Task-Relevant Units
by: AlKhamissi, Badr, et al.
Published: (2024) -
Rosetta Stone at KSAA-RD Shared Task: A Hop From Language Modeling To Word--Definition Alignment
by: ElBakry, Ahmed, et al.
Published: (2023) -
Brain-Like Language Processing via a Shallow Untrained Multihead Attention Network
by: AlKhamissi, Badr, et al.
Published: (2024) -
Inducing Dyslexia in Vision Language Models
by: Honarmand, Melika, et al.
Published: (2025) -
Investigating Cultural Alignment of Large Language Models
by: AlKhamissi, Badr, et al.
Published: (2024)