Saved in:
| Main Authors: | Surana, Risha, Ye, Qinyuan, Swayamdipta, Swabha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.10027 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Every Language Model Has a Forgery-Resistant Signature
by: Finlayson, Matthew, et al.
Published: (2025)
by: Finlayson, Matthew, et al.
Published: (2025)
BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
by: Diddee, Harshita, et al.
Published: (2026)
by: Diddee, Harshita, et al.
Published: (2026)
Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information
by: Ethayarajh, Kawin, et al.
Published: (2021)
by: Ethayarajh, Kawin, et al.
Published: (2021)
Logits of API-Protected LLMs Leak Proprietary Information
by: Finlayson, Matthew, et al.
Published: (2024)
by: Finlayson, Matthew, et al.
Published: (2024)
MixRx: Predicting Drug Combination Interactions with LLMs
by: Surana, Risha, et al.
Published: (2025)
by: Surana, Risha, et al.
Published: (2025)
CNSight: Evaluation of Clinical Note Segmentation Tools
by: Surana, Risha, et al.
Published: (2025)
by: Surana, Risha, et al.
Published: (2025)
Side-by-side Comparison Amplifies Dialect Bias in Language Models
by: Kondapally, Kritee, et al.
Published: (2026)
by: Kondapally, Kritee, et al.
Published: (2026)
Annotating FrameNet via Structure-Conditioned Language Generation
by: Cui, Xinyue, et al.
Published: (2024)
by: Cui, Xinyue, et al.
Published: (2024)
Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
by: Kulkarni, Atharva, et al.
Published: (2025)
by: Kulkarni, Atharva, et al.
Published: (2025)
Stress-Testing Long-Context Language Models with Lifelong ICL and Task Haystack
by: Xu, Xiaoyue, et al.
Published: (2024)
by: Xu, Xiaoyue, et al.
Published: (2024)
How Reliable is Language Model Micro-Benchmarking?
by: Yauney, Gregory, et al.
Published: (2025)
by: Yauney, Gregory, et al.
Published: (2025)
Compare without Despair: Reliable Preference Evaluation with Generation Separability
by: Ghosh, Sayan, et al.
Published: (2024)
by: Ghosh, Sayan, et al.
Published: (2024)
Why Fine-Tuning Encourages Hallucinations and How to Fix It
by: Kaplan, Guy, et al.
Published: (2026)
by: Kaplan, Guy, et al.
Published: (2026)
Disentangling Geometry, Performance, and Training in Language Models
by: Kulkarni, Atharva, et al.
Published: (2026)
by: Kulkarni, Atharva, et al.
Published: (2026)
GPTQT: Quantize Large Language Models Twice to Push the Efficiency
by: Guo, Yipin, et al.
Published: (2024)
by: Guo, Yipin, et al.
Published: (2024)
Design and Realization of a Benchmarking Testbed for Evaluating Autonomous Platooning Algorithms
by: Shaham, Michael, et al.
Published: (2024)
by: Shaham, Michael, et al.
Published: (2024)
Evaluation Under Imperfect Benchmarks and Ratings: A Case Study in Text Simplification
by: Liu, Joseph, et al.
Published: (2025)
by: Liu, Joseph, et al.
Published: (2025)
Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition
by: Ye, Qinyuan, et al.
Published: (2025)
by: Ye, Qinyuan, et al.
Published: (2025)
Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge
by: Cui, Xinyue, et al.
Published: (2025)
by: Cui, Xinyue, et al.
Published: (2025)
Safe Inputs but Unsafe Output: Benchmarking Cross-modality Safety Alignment of Large Vision-Language Model
by: Wang, Siyin, et al.
Published: (2024)
by: Wang, Siyin, et al.
Published: (2024)
Prompt Engineering a Prompt Engineer
by: Ye, Qinyuan, et al.
Published: (2023)
by: Ye, Qinyuan, et al.
Published: (2023)
Improving Language Model Personas via Rationalization with Psychological Scaffolds
by: Joshi, Brihi, et al.
Published: (2025)
by: Joshi, Brihi, et al.
Published: (2025)
Better Language Model Inversion by Compactly Representing Next-Token Distributions
by: Nazir, Murtaza, et al.
Published: (2025)
by: Nazir, Murtaza, et al.
Published: (2025)
Structured Program Synthesis using LLMs: Results and Insights from the IPARC Challenge
by: Surana, Shraddha, et al.
Published: (2025)
by: Surana, Shraddha, et al.
Published: (2025)
Sample, Align, Synthesize: Graph-Based Response Synthesis with ConGrs
by: Ghosh, Sayan, et al.
Published: (2025)
by: Ghosh, Sayan, et al.
Published: (2025)
Efficient Fusion and Task Guided Embedding for End-to-end Autonomous Driving
by: Guo, Yipin, et al.
Published: (2024)
by: Guo, Yipin, et al.
Published: (2024)
Command A: An Enterprise-Ready Large Language Model
by: Cohere, Team, et al.
Published: (2025)
by: Cohere, Team, et al.
Published: (2025)
The Emergence of Social Science of Large Language Models
by: Jia, Xiao, et al.
Published: (2025)
by: Jia, Xiao, et al.
Published: (2025)
Crowd-Calibrator: Can Annotator Disagreement Inform Calibration in Subjective Tasks?
by: Khurana, Urja, et al.
Published: (2024)
by: Khurana, Urja, et al.
Published: (2024)
Are We Automating the Joy Out of Work? Designing AI to Augment Work, Not Meaning
by: Ranjit, Jaspreet, et al.
Published: (2026)
by: Ranjit, Jaspreet, et al.
Published: (2026)
Generative Explanations for Program Synthesizers
by: Nazari, Amirmohammad, et al.
Published: (2024)
by: Nazari, Amirmohammad, et al.
Published: (2024)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
by: Sinha, Neelabh, et al.
Published: (2024)
by: Sinha, Neelabh, et al.
Published: (2024)
LayerTracer: A Joint Task-Particle and Vulnerable-Layer Analysis framework for Arbitrary Large Language Model Architectures
by: Wu, Yuhang, et al.
Published: (2026)
by: Wu, Yuhang, et al.
Published: (2026)
Quantifying Semantic Emergence in Language Models
by: Chen, Hang, et al.
Published: (2024)
by: Chen, Hang, et al.
Published: (2024)
The Emergence of Altruism in Large-Language-Model Agents Society
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
by: Juvekar, Kush, et al.
Published: (2025)
by: Juvekar, Kush, et al.
Published: (2025)
Street-Level AI: Are Large Language Models Ready for Real-World Judgments?
by: Pokharel, Gaurab, et al.
Published: (2025)
by: Pokharel, Gaurab, et al.
Published: (2025)
ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations
by: Joshi, Brihi, et al.
Published: (2025)
by: Joshi, Brihi, et al.
Published: (2025)
Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
by: He, Keyu, et al.
Published: (2025)
by: He, Keyu, et al.
Published: (2025)
Large Language Models as Virtual Survey Respondents: Evaluating Sociodemographic Response Generation
by: Zhao, Jianpeng, et al.
Published: (2025)
by: Zhao, Jianpeng, et al.
Published: (2025)
Similar Items
-
Every Language Model Has a Forgery-Resistant Signature
by: Finlayson, Matthew, et al.
Published: (2025) -
BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
by: Diddee, Harshita, et al.
Published: (2026) -
Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information
by: Ethayarajh, Kawin, et al.
Published: (2021) -
Logits of API-Protected LLMs Leak Proprietary Information
by: Finlayson, Matthew, et al.
Published: (2024) -
MixRx: Predicting Drug Combination Interactions with LLMs
by: Surana, Risha, et al.
Published: (2025)