Efficient and Scalable Estimation of Tool Representations in Vector Space
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Moon, Suhong, Jha, Siddharth, Erdogan, Lutfi Eren, Kim, Sehoon, Lim, Woosang, Keutzer, Kurt, Gholami, Amir |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Characterizing Prompt Compression Methods for Long Context Inference
von: Jha, Siddharth, et al.
Veröffentlicht: (2024)
von: Jha, Siddharth, et al.
Veröffentlicht: (2024)
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2025)
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2025)
TinyAgent: Function Calling at the Edge
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2024)
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2024)
Agentic Test-Time Scaling for WebAgents
von: Lee, Nicholas, et al.
Veröffentlicht: (2026)
von: Lee, Nicholas, et al.
Veröffentlicht: (2026)
Learned Best-Effort LLM Serving
von: Jha, Siddharth, et al.
Veröffentlicht: (2024)
von: Jha, Siddharth, et al.
Veröffentlicht: (2024)
An LLM Compiler for Parallel Function Calling
von: Kim, Sehoon, et al.
Veröffentlicht: (2023)
von: Kim, Sehoon, et al.
Veröffentlicht: (2023)
Stochastic Communication Avoidance for Recommendation Systems
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2024)
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2024)
SPEED: Speculative Pipelined Execution for Efficient Decoding
von: Hooper, Coleman, et al.
Veröffentlicht: (2023)
von: Hooper, Coleman, et al.
Veröffentlicht: (2023)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
von: Maheswaran, Monishwaran, et al.
Veröffentlicht: (2025)
von: Maheswaran, Monishwaran, et al.
Veröffentlicht: (2025)
Rotate, Clip, and Partition: Towards W2A4KV4 Quantization by Integrating Rotation and Learnable Non-uniform Quantizer
von: Choi, Euntae, et al.
Veröffentlicht: (2025)
von: Choi, Euntae, et al.
Veröffentlicht: (2025)
Grouped Sequency-arranged Rotation: Optimizing Rotation Transformation for Quantization for Free
von: Choi, Euntae, et al.
Veröffentlicht: (2025)
von: Choi, Euntae, et al.
Veröffentlicht: (2025)
Multipole Attention for Efficient Long Context Reasoning
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning
von: Kim, Minkyu, et al.
Veröffentlicht: (2026)
von: Kim, Minkyu, et al.
Veröffentlicht: (2026)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
Knowledge Synthesis of Photosynthesis Research Using a Large Language Model
von: Yoon, Seungri, et al.
Veröffentlicht: (2025)
von: Yoon, Seungri, et al.
Veröffentlicht: (2025)
Residual Context Diffusion Language Models
von: Hu, Yuezhou, et al.
Veröffentlicht: (2026)
von: Hu, Yuezhou, et al.
Veröffentlicht: (2026)
semantic-features: A User-Friendly Tool for Studying Contextual Word Embeddings in Interpretable Semantic Spaces
von: Ranganathan, Jwalanthi, et al.
Veröffentlicht: (2025)
von: Ranganathan, Jwalanthi, et al.
Veröffentlicht: (2025)
ETS: Efficient Tree Search for Inference-Time Scaling
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
GraphMERT: Efficient and Scalable Distillation of Reliable Knowledge Graphs from Unstructured Data
von: Belova, Margarita, et al.
Veröffentlicht: (2025)
von: Belova, Margarita, et al.
Veröffentlicht: (2025)
SqueezeLLM: Dense-and-Sparse Quantization
von: Kim, Sehoon, et al.
Veröffentlicht: (2023)
von: Kim, Sehoon, et al.
Veröffentlicht: (2023)
NeedleChain: Measuring Intact Context Comprehension Capability of Large Language Models
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2025)
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2025)
Vector-ICL: In-context Learning with Continuous Vector Representations
von: Zhuang, Yufan, et al.
Veröffentlicht: (2024)
von: Zhuang, Yufan, et al.
Veröffentlicht: (2024)
Geotokens and Geotransformers
von: Unlu, Eren
Veröffentlicht: (2024)
von: Unlu, Eren
Veröffentlicht: (2024)
LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement
von: Lee, Nicholas, et al.
Veröffentlicht: (2024)
von: Lee, Nicholas, et al.
Veröffentlicht: (2024)
Squeezed Attention: Accelerating Long Context Length LLM Inference
von: Hooper, Coleman, et al.
Veröffentlicht: (2024)
von: Hooper, Coleman, et al.
Veröffentlicht: (2024)
AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization
von: Cemri, Mert, et al.
Veröffentlicht: (2026)
von: Cemri, Mert, et al.
Veröffentlicht: (2026)
Virtual Personas for Language Models via an Anthology of Backstories
von: Moon, Suhong, et al.
Veröffentlicht: (2024)
von: Moon, Suhong, et al.
Veröffentlicht: (2024)
MacRAG: Compress, Slice, and Scale-up for Multi-Scale Adaptive Context RAG
von: Lim, Woosang, et al.
Veröffentlicht: (2025)
von: Lim, Woosang, et al.
Veröffentlicht: (2025)
The Impact of Negated Text on Hallucination with Large Language Models
von: Seo, Jaehyung, et al.
Veröffentlicht: (2025)
von: Seo, Jaehyung, et al.
Veröffentlicht: (2025)
Call for Rigor in Reporting Quality of Instruction Tuning Data
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2025)
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2025)
ToolLibGen: Scalable Automatic Tool Creation and Aggregation for LLM Reasoning
von: Yue, Murong, et al.
Veröffentlicht: (2025)
von: Yue, Murong, et al.
Veröffentlicht: (2025)
Simplex-Optimized Hybrid Ensemble for Large Language Model Text Detection Under Generative Distribution Drif
von: Kristanto, Sepyan Purnama, et al.
Veröffentlicht: (2025)
von: Kristanto, Sepyan Purnama, et al.
Veröffentlicht: (2025)
Bridging Legal Knowledge and AI: Retrieval-Augmented Generation with Vector Stores, Knowledge Graphs, and Hierarchical Non-negative Matrix Factorization
von: Barron, Ryan C., et al.
Veröffentlicht: (2025)
von: Barron, Ryan C., et al.
Veröffentlicht: (2025)
Learning Adaptive Parallel Reasoning with Language Models
von: Pan, Jiayi, et al.
Veröffentlicht: (2025)
von: Pan, Jiayi, et al.
Veröffentlicht: (2025)
Using Images to Find Context-Independent Word Representations in Vector Space
von: Kumar, Harsh
Veröffentlicht: (2024)
von: Kumar, Harsh
Veröffentlicht: (2024)
REZE: Representation Regularization for Domain-adaptive Text Embedding Pre-finetuning
von: Lee, Seungmin, et al.
Veröffentlicht: (2026)
von: Lee, Seungmin, et al.
Veröffentlicht: (2026)
RETVec: Resilient and Efficient Text Vectorizer
von: Bursztein, Elie, et al.
Veröffentlicht: (2023)
von: Bursztein, Elie, et al.
Veröffentlicht: (2023)
LegalMidm: Use-Case-Driven Legal Domain Specialization for Korean Large Language Model
von: Jang, Youngjoon, et al.
Veröffentlicht: (2026)
von: Jang, Youngjoon, et al.
Veröffentlicht: (2026)
Beyond Demonstrations: Dynamic Vector Construction from Latent Representations
von: Cai, Wang, et al.
Veröffentlicht: (2025)
von: Cai, Wang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Characterizing Prompt Compression Methods for Long Context Inference
von: Jha, Siddharth, et al.
Veröffentlicht: (2024) -
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2025) -
TinyAgent: Function Calling at the Edge
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2024) -
Agentic Test-Time Scaling for WebAgents
von: Lee, Nicholas, et al.
Veröffentlicht: (2026) -
Learned Best-Effort LLM Serving
von: Jha, Siddharth, et al.
Veröffentlicht: (2024)