Saved in:
| Main Authors: | Dalal, Uri, Segal, Meirav, Ben-Haim, Zvika, Lahav, Dan, Nevo, Omer |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.12938 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Makes an Evaluation Useful? Common Pitfalls and Best Practices
by: Gekker, Gil, et al.
Published: (2025)
by: Gekker, Gil, et al.
Published: (2025)
A Content-Based Framework for Cybersecurity Refusal Decisions in Large Language Models
by: Linder, Noa, et al.
Published: (2026)
by: Linder, Noa, et al.
Published: (2026)
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
by: Chen, Zhipeng, et al.
Published: (2025)
by: Chen, Zhipeng, et al.
Published: (2025)
Factual Inconsistency in Data-to-Text Generation Scales Exponentially with LLM Size: A Statistical Validation
by: Mahapatra, Joy, et al.
Published: (2025)
by: Mahapatra, Joy, et al.
Published: (2025)
Consistency Is the Key: Detecting Hallucinations in LLM Generated Text By Checking Inconsistencies About Key Facts
by: Gupta, Raavi, et al.
Published: (2025)
by: Gupta, Raavi, et al.
Published: (2025)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
by: Yehudai, Asaf, et al.
Published: (2025)
by: Yehudai, Asaf, et al.
Published: (2025)
Improving Quantized Model Performance in Qualitative Analysis with Multi-Pass Prompt Verification
by: Adeseye, Aisvarya, et al.
Published: (2026)
by: Adeseye, Aisvarya, et al.
Published: (2026)
JuStRank: Benchmarking LLM Judges for System Ranking
by: Gera, Ariel, et al.
Published: (2024)
by: Gera, Ariel, et al.
Published: (2024)
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
by: Hariri, Mohsen, et al.
Published: (2025)
by: Hariri, Mohsen, et al.
Published: (2025)
Deep Knowledge-Infusion For Explainable Depression Detection
by: Dalal, Sumit, et al.
Published: (2024)
by: Dalal, Sumit, et al.
Published: (2024)
Optimizing Retrieval-Augmented Generation: Analysis of Hyperparameter Impact on Performance and Efficiency
by: Ammar, Adel, et al.
Published: (2025)
by: Ammar, Adel, et al.
Published: (2025)
MPO: Boosting LLM Agents with Meta Plan Optimization
by: Xiong, Weimin, et al.
Published: (2025)
by: Xiong, Weimin, et al.
Published: (2025)
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
by: Gambardella, Andrew, et al.
Published: (2025)
by: Gambardella, Andrew, et al.
Published: (2025)
Survey on Evaluation of LLM-based Agents
by: Yehudai, Asaf, et al.
Published: (2025)
by: Yehudai, Asaf, et al.
Published: (2025)
Dimensionality Reduction in Sentence Transformer Vector Databases with Fast Fourier Transform
by: Bulgakov, Vitaly, et al.
Published: (2024)
by: Bulgakov, Vitaly, et al.
Published: (2024)
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
by: Mohamed, Anas, et al.
Published: (2025)
by: Mohamed, Anas, et al.
Published: (2025)
Knowledge Localization in Mixture-of-Experts LLMs Using Cross-Lingual Inconsistency
by: Bandarkar, Lucas, et al.
Published: (2026)
by: Bandarkar, Lucas, et al.
Published: (2026)
DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs
by: Ko, Jongwoo, et al.
Published: (2025)
by: Ko, Jongwoo, et al.
Published: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
by: Liu, Guangda, et al.
Published: (2025)
by: Liu, Guangda, et al.
Published: (2025)
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
by: Goldman, Omer, et al.
Published: (2024)
by: Goldman, Omer, et al.
Published: (2024)
Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of Perspective
by: Franzen, Daniel, et al.
Published: (2025)
by: Franzen, Daniel, et al.
Published: (2025)
Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning
by: Jung, Jaehun, et al.
Published: (2025)
by: Jung, Jaehun, et al.
Published: (2025)
Time-To-Inconsistency: A Survival Analysis of Large Language Model Robustness to Adversarial Attacks
by: Li, Yubo, et al.
Published: (2025)
by: Li, Yubo, et al.
Published: (2025)
Semantic Sensitivities and Inconsistent Predictions: Measuring the Fragility of NLI Models
by: Arakelyan, Erik, et al.
Published: (2024)
by: Arakelyan, Erik, et al.
Published: (2024)
MultiScript30k: Leveraging Multilingual Embeddings to Extend Cross Script Parallel Data
by: Driggers-Ellis, Christopher, et al.
Published: (2025)
by: Driggers-Ellis, Christopher, et al.
Published: (2025)
LegalLens: Leveraging LLMs for Legal Violation Identification in Unstructured Text
by: Bernsohn, Dor, et al.
Published: (2024)
by: Bernsohn, Dor, et al.
Published: (2024)
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
by: Du, Yihan, et al.
Published: (2024)
by: Du, Yihan, et al.
Published: (2024)
LLM Assistance for Pediatric Depression
by: Ignashina, Mariia, et al.
Published: (2025)
by: Ignashina, Mariia, et al.
Published: (2025)
Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
by: Wu, Xiaoyuan, et al.
Published: (2025)
by: Wu, Xiaoyuan, et al.
Published: (2025)
HARP: Hesitation-Aware Reframing in Transformer Inference Pass
by: Storaï, Romain, et al.
Published: (2024)
by: Storaï, Romain, et al.
Published: (2024)
Reasoning Inconsistencies and How to Mitigate Them in Deep Learning
by: Arakelyan, Erik
Published: (2025)
by: Arakelyan, Erik
Published: (2025)
Ladder: A Model-Agnostic Framework Boosting LLM-based Machine Translation to the Next Level
by: Feng, Zhaopeng, et al.
Published: (2024)
by: Feng, Zhaopeng, et al.
Published: (2024)
What's in a Name? Auditing Large Language Models for Race and Gender Bias
by: Salinas, Alejandro, et al.
Published: (2024)
by: Salinas, Alejandro, et al.
Published: (2024)
Accelerated AI Inference via Dynamic Execution Methods
by: Barad, Haim, et al.
Published: (2024)
by: Barad, Haim, et al.
Published: (2024)
Interpretable Cross-Examination Technique (ICE-T): Using highly informative features to boost LLM performance
by: Muric, Goran, et al.
Published: (2024)
by: Muric, Goran, et al.
Published: (2024)
Forma mentis networks predict creativity ratings of short texts via interpretable artificial intelligence in human and GPT-simulated raters
by: Haim, Edith, et al.
Published: (2024)
by: Haim, Edith, et al.
Published: (2024)
Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
by: Walder, Christian, et al.
Published: (2025)
by: Walder, Christian, et al.
Published: (2025)
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
Understanding the Performance and Estimating the Cost of LLM Fine-Tuning
by: Xia, Yuchen, et al.
Published: (2024)
by: Xia, Yuchen, et al.
Published: (2024)
DOMBA: Double Model Balancing for Access-Controlled Language Models via Minimum-Bounded Aggregation
by: Segal, Tom, et al.
Published: (2024)
by: Segal, Tom, et al.
Published: (2024)
Similar Items
-
What Makes an Evaluation Useful? Common Pitfalls and Best Practices
by: Gekker, Gil, et al.
Published: (2025) -
A Content-Based Framework for Cybersecurity Refusal Decisions in Large Language Models
by: Linder, Noa, et al.
Published: (2026) -
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
by: Chen, Zhipeng, et al.
Published: (2025) -
Factual Inconsistency in Data-to-Text Generation Scales Exponentially with LLM Size: A Statistical Validation
by: Mahapatra, Joy, et al.
Published: (2025) -
Consistency Is the Key: Detecting Hallucinations in LLM Generated Text By Checking Inconsistencies About Key Facts
by: Gupta, Raavi, et al.
Published: (2025)