Saved in:
| Main Authors: | Said, Muna Numan, Zaidi, Aarib, Usman, Rabia, Okon, Sonia, Medepalli, Praneeth, Zhu, Kevin, Sharma, Vasu, O'Brien, Sean |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.01430 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts
by: Gupta, Abhay, et al.
Published: (2025)
by: Gupta, Abhay, et al.
Published: (2025)
Adaptive Originality Filtering: Rejection Based Prompting and RiddleScore for Culturally Grounded Multilingual Riddle Generation
by: Le, Duy, et al.
Published: (2025)
by: Le, Duy, et al.
Published: (2025)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
by: Csizmadia, Daniel, et al.
Published: (2025)
by: Csizmadia, Daniel, et al.
Published: (2025)
Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning
by: Lu, Leo, et al.
Published: (2025)
by: Lu, Leo, et al.
Published: (2025)
SMAGDi: Socratic Multi Agent Interaction Graph Distillation for Efficient High Accuracy Reasoning
by: Aluru, Aayush, et al.
Published: (2025)
by: Aluru, Aayush, et al.
Published: (2025)
Sarc7: Evaluating Sarcasm Detection and Generation with Seven Types and Emotion-Informed Techniques
by: Xiong, Lang, et al.
Published: (2025)
by: Xiong, Lang, et al.
Published: (2025)
Pause-Tuning for Long-Context Comprehension: A Lightweight Approach to LLM Attention Recalibration
by: Begin, James, et al.
Published: (2025)
by: Begin, James, et al.
Published: (2025)
Pruning for Performance: Efficient Idiom and Metaphor Classification in Low-Resource Konkani Using mBERT
by: Do, Timothy, et al.
Published: (2025)
by: Do, Timothy, et al.
Published: (2025)
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
by: Chou, Cheng-Ting, et al.
Published: (2025)
by: Chou, Cheng-Ting, et al.
Published: (2025)
ERGO: Entropy-guided Resetting for Generation Optimization in Multi-turn Language Models
by: Khalid, Haziq Mohammad, et al.
Published: (2025)
by: Khalid, Haziq Mohammad, et al.
Published: (2025)
Universal Neurons in GPT-2: Emergence, Persistence, and Functional Impact
by: Nandan, Advey, et al.
Published: (2025)
by: Nandan, Advey, et al.
Published: (2025)
From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs
by: Yu, Stanley, et al.
Published: (2025)
by: Yu, Stanley, et al.
Published: (2025)
WOLF: Werewolf-based Observations for LLM Deception and Falsehoods
by: Agarwal, Mrinal, et al.
Published: (2025)
by: Agarwal, Mrinal, et al.
Published: (2025)
TRUTH DECAY: Quantifying Multi-Turn Sycophancy in Language Models
by: Liu, Joshua, et al.
Published: (2025)
by: Liu, Joshua, et al.
Published: (2025)
The illusion of academic freedom and the promise of the undercommons
by: Zareen Zaidi, et al.
Published: (2026)
by: Zareen Zaidi, et al.
Published: (2026)
The Geometry of Harmfulness in LLMs through Subconcept Probing
by: Shah, McNair, et al.
Published: (2025)
by: Shah, McNair, et al.
Published: (2025)
From Bias to Balance: Detecting Facial Expression Recognition Biases in Large Multimodal Foundation Models
by: Chhua, Kaylee, et al.
Published: (2024)
by: Chhua, Kaylee, et al.
Published: (2024)
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
by: Chintapatla, Ishant, et al.
Published: (2025)
by: Chintapatla, Ishant, et al.
Published: (2025)
Rosetta-PL: Propositional Logic as a Benchmark for Large Language Model Reasoning
by: Baek, Shaun, et al.
Published: (2025)
by: Baek, Shaun, et al.
Published: (2025)
Advancing Uto-Aztecan Language Technologies: A Case Study on the Endangered Comanche Language
by: C, Jesus Alvarez, et al.
Published: (2025)
by: C, Jesus Alvarez, et al.
Published: (2025)
FAIRE: Assessing Racial and Gender Bias in AI-Driven Resume Evaluations
by: Wen, Athena, et al.
Published: (2025)
by: Wen, Athena, et al.
Published: (2025)
Interpreting the Latent Structure of Operator Precedence in Language Models
by: Yugeswardeenoo, Dharunish, et al.
Published: (2025)
by: Yugeswardeenoo, Dharunish, et al.
Published: (2025)
Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks
by: Yugeswardeenoo, Dharunish, et al.
Published: (2024)
by: Yugeswardeenoo, Dharunish, et al.
Published: (2024)
MALIBU Benchmark: Multi-Agent LLM Implicit Bias Uncovered
by: Mirza, Imran, et al.
Published: (2025)
by: Mirza, Imran, et al.
Published: (2025)
Influence of Rhizophagus irregularis Inoculation on Salt Tolerance in Cucurbita maxima Duch.
by: Okon, Okon Godwin, et al.
Published: (2018)
by: Okon, Okon Godwin, et al.
Published: (2018)
Probing Audio-Generation Capabilities of Text-Based Language Models
by: Anbazhagan, Arjun Prasaath, et al.
Published: (2025)
by: Anbazhagan, Arjun Prasaath, et al.
Published: (2025)
Deconstructing FastText
by: Majumdar, Partha
Published: (2026)
by: Majumdar, Partha
Published: (2026)
Impact of Green Knowledge Sharing on the Organizational Performance of SMEs : The Mediating Role of Green Organizational Culture and Technological Innovation
by: Fernando Almeida, et al.
Published: (2026)
by: Fernando Almeida, et al.
Published: (2026)
Error Reflection Prompting: Can Large Language Models Successfully Understand Errors?
by: Li, Jason, et al.
Published: (2025)
by: Li, Jason, et al.
Published: (2025)
CLEAR: Contrasting Textual Feedback with Experts and Amateurs for Reasoning
by: Rufail, Andrew, et al.
Published: (2025)
by: Rufail, Andrew, et al.
Published: (2025)
Comparison of Radiofrequency Microneedling and Ultrasound Delivery of Plant‐Based Derived Secretory Factor (CFa1) Hair Serum for the Cosmetic Improvement of Androgenetic Alopecia
by: Lauren S. Mohan, et al.
Published: (2026)
by: Lauren S. Mohan, et al.
Published: (2026)
Rewrite-to-Rank: Optimizing Ad Visibility via Retrieval-Aware Text Rewriting
by: Ho, Chloe, et al.
Published: (2025)
by: Ho, Chloe, et al.
Published: (2025)
Flavour Deconstructing the Composite Higgs
by: Covone, Sebastiano, et al.
Published: (2024)
by: Covone, Sebastiano, et al.
Published: (2024)
Vopěnka's Principle, Maximum Deconstructibility, and singly-generated torsion classes
by: Cox, Sean
Published: (2024)
by: Cox, Sean
Published: (2024)
AAVENUE: Detecting LLM Biases on NLU Tasks in AAVE via a Novel Benchmark
by: Gupta, Abhay, et al.
Published: (2024)
by: Gupta, Abhay, et al.
Published: (2024)
Ultrafast Superconducting Qubit Readout with the Quarton Coupler
by: Ye, Yufeng, et al.
Published: (2024)
by: Ye, Yufeng, et al.
Published: (2024)
El problema de la medición en mecánica cuántica
by: E. Okon
Published: (2014)
by: E. Okon
Published: (2014)
Reassessing the strength of a class of Wigner's friend no-go theorems
by: Okon, E.
Published: (2022)
by: Okon, E.
Published: (2022)
Deconstructive Composite Dark Matter Detection
by: Boukhtouchen, Yilda, et al.
Published: (2025)
by: Boukhtouchen, Yilda, et al.
Published: (2025)
Encoding Inequity: Examining Demographic Bias in LLM-Driven Robot Caregiving
by: Korpan, Raj
Published: (2025)
by: Korpan, Raj
Published: (2025)
Similar Items
-
NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts
by: Gupta, Abhay, et al.
Published: (2025) -
Adaptive Originality Filtering: Rejection Based Prompting and RiddleScore for Culturally Grounded Multilingual Riddle Generation
by: Le, Duy, et al.
Published: (2025) -
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
by: Csizmadia, Daniel, et al.
Published: (2025) -
Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning
by: Lu, Leo, et al.
Published: (2025) -
SMAGDi: Socratic Multi Agent Interaction Graph Distillation for Efficient High Accuracy Reasoning
by: Aluru, Aayush, et al.
Published: (2025)