Saved in:
| Main Authors: | Liu, Haoran, Tahmasbi, Amir, Haque, Ehtesham Sam, Jain, Purak |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.17863 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Building Blocks to Planning: Multi-Step Spatial Reasoning in LLMs with Reinforcement Learning
by: Tahmasbi, Amir, et al.
Published: (2025)
by: Tahmasbi, Amir, et al.
Published: (2025)
Scaling Laws for Discriminative Classification in Large Language Models
by: Wyatte, Dean, et al.
Published: (2024)
by: Wyatte, Dean, et al.
Published: (2024)
SERP Interference Network and Its Applications in Search Advertising
by: Jain, Purak, et al.
Published: (2025)
by: Jain, Purak, et al.
Published: (2025)
Market-Bench: Evaluating Large Language Models on Introductory Quantitative Trading and Market Dynamics
by: Srivastava, Abhay, et al.
Published: (2025)
by: Srivastava, Abhay, et al.
Published: (2025)
SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models
by: Gu, Yiyang, et al.
Published: (2026)
by: Gu, Yiyang, et al.
Published: (2026)
Catastrophic Forgetting in LLMs: A Comparative Analysis Across Language Tasks
by: Haque, Naimul
Published: (2025)
by: Haque, Naimul
Published: (2025)
Evaluating Alignment of Behavioral Dispositions in LLMs
by: Taubenfeld, Amir, et al.
Published: (2026)
by: Taubenfeld, Amir, et al.
Published: (2026)
Self-Improving Customer Review Response Generation Based on LLMs
by: Azov, Guy, et al.
Published: (2024)
by: Azov, Guy, et al.
Published: (2024)
IQ Test for LLMs: An Evaluation Framework for Uncovering Core Skills in LLMs
by: Maimon, Aviya, et al.
Published: (2025)
by: Maimon, Aviya, et al.
Published: (2025)
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
by: Xie, Yuxuan, et al.
Published: (2024)
by: Xie, Yuxuan, et al.
Published: (2024)
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG
by: Singh, Aditi, et al.
Published: (2025)
by: Singh, Aditi, et al.
Published: (2025)
Large Language Models as Pokémon Battle Agents: Strategic Play and Content Generation
by: Jain, Daksh, et al.
Published: (2025)
by: Jain, Daksh, et al.
Published: (2025)
Cross-Asset Risk Management: Integrating LLMs for Real-Time Monitoring of Equity, Fixed Income, and Currency Markets
by: Yang, Jie, et al.
Published: (2025)
by: Yang, Jie, et al.
Published: (2025)
Digital Diagnostics: The Potential Of Large Language Models In Recognizing Symptoms Of Common Illnesses
by: Gupta, Gaurav Kumar, et al.
Published: (2024)
by: Gupta, Gaurav Kumar, et al.
Published: (2024)
Trustful LLMs: Customizing and Grounding Text Generation with Knowledge Bases and Dual Decoders
by: Zhu, Xiaofeng, et al.
Published: (2024)
by: Zhu, Xiaofeng, et al.
Published: (2024)
What's Taboo for You? - An Empirical Evaluation of LLMs Behavior Toward Sensitive Content
by: Ferrara, Alfio, et al.
Published: (2025)
by: Ferrara, Alfio, et al.
Published: (2025)
ToPSen: Task-Oriented Priming and Sensory Alignment for Comparing Coding Strategies Between Sighted and Blind Programmers
by: Ehtesham-Ul-Haque, Md, et al.
Published: (2025)
by: Ehtesham-Ul-Haque, Md, et al.
Published: (2025)
VoiceAlign: A Shimming Layer for Enhancing the Usability of Legacy Voice User Interface Systems
by: Ehtesham-Ul-Haque, Md, et al.
Published: (2026)
by: Ehtesham-Ul-Haque, Md, et al.
Published: (2026)
Architectural Flaw Detection in Civil Engineering Using GPT-4
by: Kumar, Saket, et al.
Published: (2024)
by: Kumar, Saket, et al.
Published: (2024)
CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs
by: Wu, Chengwei, et al.
Published: (2025)
by: Wu, Chengwei, et al.
Published: (2025)
Scaling laws for nonlinear dynamical models of articulatory control
by: Kirkham, Sam
Published: (2024)
by: Kirkham, Sam
Published: (2024)
MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
Evaluating, Synthesizing, and Enhancing for Customer Support Conversation
by: Zhu, Jie, et al.
Published: (2025)
by: Zhu, Jie, et al.
Published: (2025)
Intent Detection in the Age of LLMs
by: Arora, Gaurav, et al.
Published: (2024)
by: Arora, Gaurav, et al.
Published: (2024)
Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs
by: Kumar, Divyanshu, et al.
Published: (2024)
by: Kumar, Divyanshu, et al.
Published: (2024)
Harnessing LLMs for Educational Content-Driven Italian Crossword Generation
by: Zeinalipour, Kamyar, et al.
Published: (2024)
by: Zeinalipour, Kamyar, et al.
Published: (2024)
Ideology-Based LLMs for Content Moderation
by: Civelli, Stefano, et al.
Published: (2025)
by: Civelli, Stefano, et al.
Published: (2025)
Scheherazade: Evaluating Chain-of-Thought Math Reasoning in LLMs with Chain-of-Problems
by: Miner, Stephen, et al.
Published: (2024)
by: Miner, Stephen, et al.
Published: (2024)
LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling
by: Zheng, Tong, et al.
Published: (2026)
by: Zheng, Tong, et al.
Published: (2026)
From Synthetic to Native: Benchmarking Multilingual Intent Classification in Logistics Customer Service
by: He, Haoyu, et al.
Published: (2026)
by: He, Haoyu, et al.
Published: (2026)
Evaluating the Diversity and Quality of LLM Generated Content
by: Shypula, Alexander, et al.
Published: (2025)
by: Shypula, Alexander, et al.
Published: (2025)
UniSparse: An Intermediate Language for General Sparse Format Customization
by: Liu, Jie, et al.
Published: (2024)
by: Liu, Jie, et al.
Published: (2024)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
by: Wang, Yuxia, et al.
Published: (2024)
by: Wang, Yuxia, et al.
Published: (2024)
LLM-Assisted Question-Answering on Technical Documents Using Structured Data-Aware Retrieval Augmented Generation
by: Sobhan, Shadman, et al.
Published: (2025)
by: Sobhan, Shadman, et al.
Published: (2025)
PersonaBOT: Bringing Customer Personas to Life with LLMs and RAG
by: Rizwan, Muhammed, et al.
Published: (2025)
by: Rizwan, Muhammed, et al.
Published: (2025)
Systematic Evaluation of the Quality of Synthetic Clinical Notes Rephrased by LLMs at Million-Note Scale
by: Liu, Jinghui, et al.
Published: (2026)
by: Liu, Jinghui, et al.
Published: (2026)
Customizing Large Language Model Generation Style using Parameter-Efficient Finetuning
by: Liu, Xinyue, et al.
Published: (2024)
by: Liu, Xinyue, et al.
Published: (2024)
LLMs as Models for Analogical Reasoning
by: Musker, Sam, et al.
Published: (2024)
by: Musker, Sam, et al.
Published: (2024)
Personalized Causal Graph Reasoning for LLMs: An Implementation for Dietary Recommendations
by: Yang, Zhongqi, et al.
Published: (2025)
by: Yang, Zhongqi, et al.
Published: (2025)
Similar Items
-
From Building Blocks to Planning: Multi-Step Spatial Reasoning in LLMs with Reinforcement Learning
by: Tahmasbi, Amir, et al.
Published: (2025) -
Scaling Laws for Discriminative Classification in Large Language Models
by: Wyatte, Dean, et al.
Published: (2024) -
SERP Interference Network and Its Applications in Search Advertising
by: Jain, Purak, et al.
Published: (2025) -
Market-Bench: Evaluating Large Language Models on Introductory Quantitative Trading and Market Dynamics
by: Srivastava, Abhay, et al.
Published: (2025) -
SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models
by: Gu, Yiyang, et al.
Published: (2026)