Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yixin, Xiong, Ying, Wu, Shangyu, Cui, Yufei, Liu, Xue, Guan, Nan, Xue, Chun Jason |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReFilter: Improving Robustness of Retrieval-Augmented Generation via Gated Filter
by: Chen, Yixin, et al.
Published: (2026)
by: Chen, Yixin, et al.
Published: (2026)
RAEE: A Robust Retrieval-Augmented Early Exit Framework for Efficient Inference
by: Huang, Lianming, et al.
Published: (2024)
by: Huang, Lianming, et al.
Published: (2024)
ReFusion: Improving Natural Language Understanding with Computation-Efficient Retrieval Representation Fusion
by: Wu, Shangyu, et al.
Published: (2024)
by: Wu, Shangyu, et al.
Published: (2024)
Retrieval-Augmented Generation for Natural Language Processing: A Survey
by: Wu, Shangyu, et al.
Published: (2024)
by: Wu, Shangyu, et al.
Published: (2024)
EvoP: Robust LLM Inference via Evolutionary Pruning
by: Wu, Shangyu, et al.
Published: (2025)
by: Wu, Shangyu, et al.
Published: (2025)
A$^2$ATS: Retrieval-Based KV Cache Reduction via Windowed Rotary Position Embedding and Query-Aware Vector Quantization
by: He, Junhui, et al.
Published: (2025)
by: He, Junhui, et al.
Published: (2025)
CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
by: He, Junhui, et al.
Published: (2024)
by: He, Junhui, et al.
Published: (2024)
BAHOP: Similarity-based Basin Hopping for A fast hyper-parameter search in WSI classification
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
GeneQuery: A General QA-based Framework for Spatial Gene Expression Predictions from Histology Images
by: Xiong, Ying, et al.
Published: (2024)
by: Xiong, Ying, et al.
Published: (2024)
Nav-EE: Navigation-Guided Early Exiting for Efficient Vision-Language Models in Autonomous Driving
by: Hu, Haibo, et al.
Published: (2025)
by: Hu, Haibo, et al.
Published: (2025)
AD-EE: Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving
by: Huang, Lianming, et al.
Published: (2025)
by: Huang, Lianming, et al.
Published: (2025)
FlexInfer: Breaking Memory Constraint via Flexible and Efficient Offloading for On-Device LLM Inference
by: Du, Hongchao, et al.
Published: (2025)
by: Du, Hongchao, et al.
Published: (2025)
IHC Matters: Incorporating IHC analysis to H&E Whole Slide Image Analysis for Improved Cancer Grading via Two-stage Multimodal Bilinear Pooling Fusion
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
by: Zou, Jing, et al.
Published: (2026)
by: Zou, Jing, et al.
Published: (2026)
On the Compressibility of Quantized Large Language Models
by: Mao, Yu, et al.
Published: (2024)
by: Mao, Yu, et al.
Published: (2024)
Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
by: Srinivasan, Tejas, et al.
Published: (2024)
by: Srinivasan, Tejas, et al.
Published: (2024)
Intuitive or Dependent? Investigating LLMs' Behavior Style to Conflicting Prompts
by: Ying, Jiahao, et al.
Published: (2023)
by: Ying, Jiahao, et al.
Published: (2023)
When Compression Meets Model Compression: Memory-Efficient Double Compression for Large Language Models
by: Wang, Weilan, et al.
Published: (2025)
by: Wang, Weilan, et al.
Published: (2025)
Tool Calling: Enhancing Medication Consultation via Retrieval-Augmented Large Language Models
by: Huang, Zhongzhen, et al.
Published: (2024)
by: Huang, Zhongzhen, et al.
Published: (2024)
Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation
by: Marketsmüller, Michael, et al.
Published: (2026)
by: Marketsmüller, Michael, et al.
Published: (2026)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
by: Mao, Yu, et al.
Published: (2025)
by: Mao, Yu, et al.
Published: (2025)
Beyond Semantic Similarity: A Two-Phase Non-Parametric Retrieval Workflow for Corporate Credit Underwriting
by: Junjia, Linus Ng, et al.
Published: (2026)
by: Junjia, Linus Ng, et al.
Published: (2026)
API Pack: A Massive Multi-Programming Language Dataset for API Call Generation
by: Guo, Zhen, et al.
Published: (2024)
by: Guo, Zhen, et al.
Published: (2024)
From Detection to Mitigation: Addressing Gender Bias in Chinese Texts via Efficient Tuning and Voting-Based Rebalancing
by: Wu, Chengyan, et al.
Published: (2025)
by: Wu, Chengyan, et al.
Published: (2025)
Cache & Distil: Optimising API Calls to Large Language Models
by: Ramírez, Guillem, et al.
Published: (2023)
by: Ramírez, Guillem, et al.
Published: (2023)
AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls
by: Du, Yu, et al.
Published: (2024)
by: Du, Yu, et al.
Published: (2024)
Transforming Questions and Documents for Semantically Aligned Retrieval-Augmented Generation
by: Lee, Seokgi
Published: (2025)
by: Lee, Seokgi
Published: (2025)
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
by: Basu, Kinjal, et al.
Published: (2024)
by: Basu, Kinjal, et al.
Published: (2024)
Beyond Semantic Entropy: Boosting LLM Uncertainty Quantification with Pairwise Semantic Similarity
by: Nguyen, Dang, et al.
Published: (2025)
by: Nguyen, Dang, et al.
Published: (2025)
Beyond Surface Similarity: Detecting Subtle Semantic Shifts in Financial Narratives
by: Liu, Jiaxin, et al.
Published: (2024)
by: Liu, Jiaxin, et al.
Published: (2024)
Are ELECTRA's Sentence Embeddings Beyond Repair? The Case of Semantic Textual Similarity
by: Rep, Ivan, et al.
Published: (2024)
by: Rep, Ivan, et al.
Published: (2024)
Beyond Semantic Relevance: Counterfactual Risk Minimization for Robust Retrieval-Augmented Generation
by: Liu, Peiyang, et al.
Published: (2026)
by: Liu, Peiyang, et al.
Published: (2026)
Evaluating LLMs on Sequential API Call Through Automated Test Generation
by: Huang, Yuheng, et al.
Published: (2025)
by: Huang, Yuheng, et al.
Published: (2025)
Beyond Retrieval: Ensembling Cross-Encoders and GPT Rerankers with LLMs for Biomedical QA
by: Verma, Shashank, et al.
Published: (2025)
by: Verma, Shashank, et al.
Published: (2025)
Semantic Convergence: Harmonizing Recommender Systems via Two-Stage Alignment and Behavioral Semantic Tokenization
by: Li, Guanghan, et al.
Published: (2024)
by: Li, Guanghan, et al.
Published: (2024)
Beyond Topical Similarity: Contrastive Evidence Retrieval with Interpretable Attention Alignment in RAG
by: Vargas, Francielle, et al.
Published: (2026)
by: Vargas, Francielle, et al.
Published: (2026)
Association Is Not Similarity: Learning Corpus-Specific Associations for Multi-Hop Retrieval
by: Dury, Jason
Published: (2026)
by: Dury, Jason
Published: (2026)
Beyond Browsing: API-Based Web Agents
by: Song, Yueqi, et al.
Published: (2024)
by: Song, Yueqi, et al.
Published: (2024)
Tool Retrieval Bridge: Aligning Vague Instructions with Retriever Preferences via Bridge Model
by: Chen, Kunfeng, et al.
Published: (2026)
by: Chen, Kunfeng, et al.
Published: (2026)
Direct Semantic Communication Between Large Language Models via Vector Translation
by: Yang, Fu-Chun, et al.
Published: (2025)
by: Yang, Fu-Chun, et al.
Published: (2025)
Similar Items
-
ReFilter: Improving Robustness of Retrieval-Augmented Generation via Gated Filter
by: Chen, Yixin, et al.
Published: (2026) -
RAEE: A Robust Retrieval-Augmented Early Exit Framework for Efficient Inference
by: Huang, Lianming, et al.
Published: (2024) -
ReFusion: Improving Natural Language Understanding with Computation-Efficient Retrieval Representation Fusion
by: Wu, Shangyu, et al.
Published: (2024) -
Retrieval-Augmented Generation for Natural Language Processing: A Survey
by: Wu, Shangyu, et al.
Published: (2024) -
EvoP: Robust LLM Inference via Evolutionary Pruning
by: Wu, Shangyu, et al.
Published: (2025)