Saved in:
| Main Authors: | Gul, Haji, Naim, Abul Ghani, Bhat, Ajaz Ahmad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2508.15357 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Knowledge Graph Complexity via Semantic, Spectral, and Structural Metrics for Link Prediction
by: Gul, Haji, et al.
Published: (2025)
by: Gul, Haji, et al.
Published: (2025)
MuCo-KGC: Multi-Context-Aware Knowledge Graph Completion
by: Gul, Haji, et al.
Published: (2025)
by: Gul, Haji, et al.
Published: (2025)
A Contextualized BERT model for Knowledge Graph Completion
by: Gul, Haji, et al.
Published: (2024)
by: Gul, Haji, et al.
Published: (2024)
MuCoS: Efficient Drug Target Discovery via Multi Context Aware Sampling in Knowledge Graphs
by: Gul, Haji, et al.
Published: (2025)
by: Gul, Haji, et al.
Published: (2025)
Evaluating Cumulative Spectral Gradient as a Complexity Measure
by: Gul, Haji, et al.
Published: (2025)
by: Gul, Haji, et al.
Published: (2025)
MuCoS: Efficient Drug-Target Prediction through Multi-Context-Aware Sampling
by: Gul, Haji, et al.
Published: (2025)
by: Gul, Haji, et al.
Published: (2025)
LFED: A Literary Fiction Evaluation Dataset for Large Language Models
by: Yu, Linhao, et al.
Published: (2024)
by: Yu, Linhao, et al.
Published: (2024)
Systematic Evaluation of Optimization Techniques for Long-Context Language Models
by: Ahmed, Ammar, et al.
Published: (2025)
by: Ahmed, Ammar, et al.
Published: (2025)
Layer Importance and Hallucination Analysis in Large Language Models via Enhanced Activation Variance-Sparsity
by: Song, Zichen, et al.
Published: (2024)
by: Song, Zichen, et al.
Published: (2024)
Metrics and evaluations for computational and sustainable AI efficiency
by: Liu, Hongyuan, et al.
Published: (2025)
by: Liu, Hongyuan, et al.
Published: (2025)
Green AI: Exploring Carbon Footprints, Mitigation Strategies, and Trade Offs in Large Language Model Training
by: Liu, Vivian, et al.
Published: (2024)
by: Liu, Vivian, et al.
Published: (2024)
SimLens for Early Exit in Large Language Models: Eliciting Accurate Latent Predictions with One More Token
by: Ma, Ming, et al.
Published: (2025)
by: Ma, Ming, et al.
Published: (2025)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
by: Chen, Feiyang, et al.
Published: (2025)
by: Chen, Feiyang, et al.
Published: (2025)
SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version
by: Pham, Nghiem Thanh, et al.
Published: (2025)
by: Pham, Nghiem Thanh, et al.
Published: (2025)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
by: Chiang, Hung-Yueh, et al.
Published: (2025)
by: Chiang, Hung-Yueh, et al.
Published: (2025)
Iterative Layer Pruning for Efficient Translation Inference
by: Moslem, Yasmin, et al.
Published: (2025)
by: Moslem, Yasmin, et al.
Published: (2025)
L1RA: Dynamic Rank Assignment in LoRA Fine-Tuning
by: Singh, Raul, et al.
Published: (2025)
by: Singh, Raul, et al.
Published: (2025)
Beyond Tokens: Semantic-Aware Speculative Decoding for Efficient Inference by Probing Internal States
by: Dong, Ximing, et al.
Published: (2026)
by: Dong, Ximing, et al.
Published: (2026)
Evaluating the Efficacy of Foundational Models: Advancing Benchmarking Practices to Enhance Fine-Tuning Decision-Making
by: Amujo, Oluyemi Enoch, et al.
Published: (2024)
by: Amujo, Oluyemi Enoch, et al.
Published: (2024)
Priority Sampling of Large Language Models for Compilers
by: Grubisic, Dejan, et al.
Published: (2024)
by: Grubisic, Dejan, et al.
Published: (2024)
Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey
by: Moslem, Yasmin, et al.
Published: (2026)
by: Moslem, Yasmin, et al.
Published: (2026)
Model Compression and Efficient Inference for Large Language Models: A Survey
by: Wang, Wenxiao, et al.
Published: (2024)
by: Wang, Wenxiao, et al.
Published: (2024)
A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
by: Agullo, Ferran, et al.
Published: (2025)
by: Agullo, Ferran, et al.
Published: (2025)
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
by: Liu, Zirui, et al.
Published: (2024)
by: Liu, Zirui, et al.
Published: (2024)
ModeSwitch-LLM: A Lightweight Phase-Aware Controller for Cross-Mode LLM Inference on a Single GPU
by: Sunesh, Aman, et al.
Published: (2026)
by: Sunesh, Aman, et al.
Published: (2026)
Rule-Based Graph Programs Matching the Time Complexity of Imperative Algorithms
by: Alaoui, Ziad Ismaili, et al.
Published: (2025)
by: Alaoui, Ziad Ismaili, et al.
Published: (2025)
SuperCoder: Assembly Program Superoptimization with Large Language Models
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Harnessing Chain-of-Thought Metadata for Task Routing and Adversarial Prompt Detection
by: Marinelli, Ryan, et al.
Published: (2025)
by: Marinelli, Ryan, et al.
Published: (2025)
DeepSeek-R1 Outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in Bilingual Complex Ophthalmology Reasoning
by: Xu, Pusheng, et al.
Published: (2025)
by: Xu, Pusheng, et al.
Published: (2025)
TeleEval-OS: Performance evaluations of large language models for operations scheduling
by: Wang, Yanyan, et al.
Published: (2025)
by: Wang, Yanyan, et al.
Published: (2025)
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
by: Liu, Minghui, et al.
Published: (2025)
by: Liu, Minghui, et al.
Published: (2025)
Block Sparse Flash Attention
by: Ohayon, Daniel, et al.
Published: (2025)
by: Ohayon, Daniel, et al.
Published: (2025)
Evaluating Compiler Optimization Impacts on zkVM Performance
by: Gassmann, Thomas, et al.
Published: (2025)
by: Gassmann, Thomas, et al.
Published: (2025)
Embedding Retrofitting: Data Engineering for better RAG
by: Sharma, Anantha
Published: (2026)
by: Sharma, Anantha
Published: (2026)
DyLLM: Efficient Diffusion LLM Inference via Saliency-based Token Selection and Partial Attention
by: Lee, Younjoo, et al.
Published: (2026)
by: Lee, Younjoo, et al.
Published: (2026)
FLuRKA: Fast and accurate unified Low-Rank & Kernel Attention
by: Gupta, Ahan, et al.
Published: (2023)
by: Gupta, Ahan, et al.
Published: (2023)
Performance Characterization of Expert Router for Scalable LLM Inference
by: Pichlmeier, Josef, et al.
Published: (2024)
by: Pichlmeier, Josef, et al.
Published: (2024)
Data Efficacy for Language Model Training
by: Dai, Yalun, et al.
Published: (2025)
by: Dai, Yalun, et al.
Published: (2025)
Investigating Execution-Aware Language Models for Code Optimization
by: Di Menna, Federico, et al.
Published: (2025)
by: Di Menna, Federico, et al.
Published: (2025)
Runtime Repeated Recursion Unfolding in CHR: A Just-In-Time Online Program Optimization Strategy That Can Achieve Super-Linear Speedup
by: Fruehwirth, Thom
Published: (2023)
by: Fruehwirth, Thom
Published: (2023)
Similar Items
-
Evaluating Knowledge Graph Complexity via Semantic, Spectral, and Structural Metrics for Link Prediction
by: Gul, Haji, et al.
Published: (2025) -
MuCo-KGC: Multi-Context-Aware Knowledge Graph Completion
by: Gul, Haji, et al.
Published: (2025) -
A Contextualized BERT model for Knowledge Graph Completion
by: Gul, Haji, et al.
Published: (2024) -
MuCoS: Efficient Drug Target Discovery via Multi Context Aware Sampling in Knowledge Graphs
by: Gul, Haji, et al.
Published: (2025) -
Evaluating Cumulative Spectral Gradient as a Complexity Measure
by: Gul, Haji, et al.
Published: (2025)