Saved in:
| Main Author: | Sahasrabudhe, Mihir |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.19997 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reversal Invariance in Autoregressive Language Models
by: Sahasrabudhe, Mihir
Published: (2025)
by: Sahasrabudhe, Mihir
Published: (2025)
Marginals Before Conditionals
by: Sahasrabudhe, Mihir
Published: (2026)
by: Sahasrabudhe, Mihir
Published: (2026)
Support-Contra Asymmetry in LLM Explanations
by: Patil, Avinash
Published: (2025)
by: Patil, Avinash
Published: (2025)
Towards Autoformalization of LLM-generated Outputs for Requirement Verification
by: Gupte, Mihir, et al.
Published: (2025)
by: Gupte, Mihir, et al.
Published: (2025)
MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use
by: Wu, Zijian, et al.
Published: (2025)
by: Wu, Zijian, et al.
Published: (2025)
Testing the Limits of Truth Directions in LLMs
by: Poulis, Angelos, et al.
Published: (2026)
by: Poulis, Angelos, et al.
Published: (2026)
What Works for 'Lost-in-the-Middle' in LLMs? A Study on GM-Extract and Mitigations
by: Gupte, Mihir, et al.
Published: (2025)
by: Gupte, Mihir, et al.
Published: (2025)
Indic-TunedLens: Interpreting Multilingual Models in Indian Languages
by: Panchal, Mihir, et al.
Published: (2026)
by: Panchal, Mihir, et al.
Published: (2026)
Enhancing Transformer-Based Rerankers with Synthetic Data and LLM-Based Supervision
by: Peshevski, Dimitar, et al.
Published: (2025)
by: Peshevski, Dimitar, et al.
Published: (2025)
Stress-Testing Model Specs Reveals Character Differences among Language Models
by: Zhang, Jifan, et al.
Published: (2025)
by: Zhang, Jifan, et al.
Published: (2025)
Optimsyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation
by: Fan, Zhiting, et al.
Published: (2026)
by: Fan, Zhiting, et al.
Published: (2026)
A Dual-Directional Context-Aware Test-Time Learning for Text Classification
by: Xu, Dong, et al.
Published: (2025)
by: Xu, Dong, et al.
Published: (2025)
Token-level Direct Preference Optimization
by: Zeng, Yongcheng, et al.
Published: (2024)
by: Zeng, Yongcheng, et al.
Published: (2024)
LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
by: Yin, Ming, et al.
Published: (2025)
by: Yin, Ming, et al.
Published: (2025)
Communication and Verification in LLM Agents towards Collaboration under Information Asymmetry
by: Peng, Run, et al.
Published: (2025)
by: Peng, Run, et al.
Published: (2025)
A Fixed-Budget, Cluster-Aware Standard for LLM-as-a-Judge Evaluation: A Multi-Hop RAG Stress Test
by: Sartori, Camilo Chacón, et al.
Published: (2026)
by: Sartori, Camilo Chacón, et al.
Published: (2026)
The Fellowship of the LLMs: Multi-Model Workflows for Synthetic Preference Optimization Dataset Generation
by: Arif, Samee, et al.
Published: (2024)
by: Arif, Samee, et al.
Published: (2024)
Is Implicit Knowledge Enough for LLMs? A RAG Approach for Tree-based Structures
by: Gupte, Mihir, et al.
Published: (2025)
by: Gupte, Mihir, et al.
Published: (2025)
Stress Testing Factual Consistency Metrics for Long-Document Summarization
by: Mujahid, Zain Muhammad, et al.
Published: (2025)
by: Mujahid, Zain Muhammad, et al.
Published: (2025)
Distributed Partial Information Puzzles: Examining Common Ground Construction Under Epistemic Asymmetry
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
Compound AI Systems Optimization: A Survey of Methods, Challenges, and Future Directions
by: Lee, Yu-Ang, et al.
Published: (2025)
by: Lee, Yu-Ang, et al.
Published: (2025)
Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark
by: Gupta, Himanshu, et al.
Published: (2024)
by: Gupta, Himanshu, et al.
Published: (2024)
SDPO: Segment-Level Direct Preference Optimization for Social Agents
by: Kong, Aobo, et al.
Published: (2025)
by: Kong, Aobo, et al.
Published: (2025)
The Power of Power Law: Asymmetry Enables Compositional Reasoning
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models
by: Wang, Wentian, et al.
Published: (2024)
by: Wang, Wentian, et al.
Published: (2024)
Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient Transformers
by: Tang, Zecheng, et al.
Published: (2026)
by: Tang, Zecheng, et al.
Published: (2026)
Filtered Direct Preference Optimization
by: Morimura, Tetsuro, et al.
Published: (2024)
by: Morimura, Tetsuro, et al.
Published: (2024)
Direct Preference Optimization with an Offset
by: Amini, Afra, et al.
Published: (2024)
by: Amini, Afra, et al.
Published: (2024)
Characterizing the Robustness of Black-Box LLM Planners Under Perturbed Observations with Adaptive Stress Testing
by: Chakraborty, Neeloy, et al.
Published: (2025)
by: Chakraborty, Neeloy, et al.
Published: (2025)
Synthetic4Health: Generating Annotated Synthetic Clinical Letters
by: Ren, Libo, et al.
Published: (2024)
by: Ren, Libo, et al.
Published: (2024)
Knowledge Editing in Language Models via Adapted Direct Preference Optimization
by: Rozner, Amit, et al.
Published: (2024)
by: Rozner, Amit, et al.
Published: (2024)
Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models
by: Luo, Man, et al.
Published: (2023)
by: Luo, Man, et al.
Published: (2023)
Stress-Testing Long-Context Language Models with Lifelong ICL and Task Haystack
by: Xu, Xiaoyue, et al.
Published: (2024)
by: Xu, Xiaoyue, et al.
Published: (2024)
Towards Enhancing Coherence in Extractive Summarization: Dataset and Experiments with LLMs
by: Parmar, Mihir, et al.
Published: (2024)
by: Parmar, Mihir, et al.
Published: (2024)
Humane Speech Synthesis through Zero-Shot Emotion and Disfluency Generation
by: Chaudhury, Rohan, et al.
Published: (2024)
by: Chaudhury, Rohan, et al.
Published: (2024)
Synthetic bootstrapped pretraining
by: Yang, Zitong, et al.
Published: (2025)
by: Yang, Zitong, et al.
Published: (2025)
Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values
by: Zhang, Hongbo, et al.
Published: (2025)
by: Zhang, Hongbo, et al.
Published: (2025)
DelvePO: Direction-Guided Self-Evolving Framework for Flexible Prompt Optimization
by: Tao, Tao, et al.
Published: (2025)
by: Tao, Tao, et al.
Published: (2025)
2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision
by: Li, Shilong, et al.
Published: (2024)
by: Li, Shilong, et al.
Published: (2024)
StressTransfer: Stress-Aware Speech-to-Speech Translation with Emphasis Preservation
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
Similar Items
-
Reversal Invariance in Autoregressive Language Models
by: Sahasrabudhe, Mihir
Published: (2025) -
Marginals Before Conditionals
by: Sahasrabudhe, Mihir
Published: (2026) -
Support-Contra Asymmetry in LLM Explanations
by: Patil, Avinash
Published: (2025) -
Towards Autoformalization of LLM-generated Outputs for Requirement Verification
by: Gupte, Mihir, et al.
Published: (2025) -
MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use
by: Wu, Zijian, et al.
Published: (2025)