Efficient LLM-based Advertising via Model Compression and Parallel Verification
Fuente:
arXiv
Salvato in:
| Autori principali: | Dong, Wenxin, Gao, Chang, Yu, Guanghui, Jiao, Xuewu, Hu, Mingqing, Fu, Qiang, Xu, Peng, Wei, Penghui, Xu, Hui, Xing, Yue, Li, Shuanglong, Liu, Lin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search for LLM Inference
di: Dong, Wenxin, et al.
Pubblicazione: (2026)
di: Dong, Wenxin, et al.
Pubblicazione: (2026)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025)
di: Peters, Sydney, et al.
Pubblicazione: (2025)
SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning
di: Chang, Edward Y., et al.
Pubblicazione: (2025)
di: Chang, Edward Y., et al.
Pubblicazione: (2025)
Active Context Compression: Autonomous Memory Management in LLM Agents
di: Verma, Nikhil
Pubblicazione: (2026)
di: Verma, Nikhil
Pubblicazione: (2026)
Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment
di: Chang, Edward Y.
Pubblicazione: (2026)
di: Chang, Edward Y.
Pubblicazione: (2026)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
di: Souza, Débora, et al.
Pubblicazione: (2026)
di: Souza, Débora, et al.
Pubblicazione: (2026)
LLM-Based SQL Generation: Prompting, Self-Refinement, and Adaptive Weighted Majority Voting
di: Yang, Yu-Jie, et al.
Pubblicazione: (2026)
di: Yang, Yu-Jie, et al.
Pubblicazione: (2026)
On the Challenges of Creating Datasets for Analyzing Commercial Sex Advertisements to Assess Human Trafficking Risk and Organized Activity
di: Rivas, Pablo, et al.
Pubblicazione: (2024)
di: Rivas, Pablo, et al.
Pubblicazione: (2024)
EVINCE: Optimizing Multi-LLM Dialogues Using Conditional Statistics and Information Theory
di: Chang, Edward Y.
Pubblicazione: (2024)
di: Chang, Edward Y.
Pubblicazione: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
LLM-based Automated Theorem Proving Hinges on Scalable Synthetic Data Generation
di: Lai, Junyu, et al.
Pubblicazione: (2025)
di: Lai, Junyu, et al.
Pubblicazione: (2025)
ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning
di: Chang, Edward Y., et al.
Pubblicazione: (2025)
di: Chang, Edward Y., et al.
Pubblicazione: (2025)
From Classification to Ranking: Enhancing LLM Reasoning Capabilities for MBTI Personality Detection
di: Cao, Yuan, et al.
Pubblicazione: (2026)
di: Cao, Yuan, et al.
Pubblicazione: (2026)
Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
di: Johnson, Warren
Pubblicazione: (2026)
di: Johnson, Warren
Pubblicazione: (2026)
PLM: Efficient Peripheral Language Models Hardware-Co-Designed for Ubiquitous Computing
di: Deng, Cheng, et al.
Pubblicazione: (2025)
di: Deng, Cheng, et al.
Pubblicazione: (2025)
LLM Vocabulary Compression for Low-Compute Environments
di: Vennam, Sreeram, et al.
Pubblicazione: (2024)
di: Vennam, Sreeram, et al.
Pubblicazione: (2024)
ML-Promise: A Multilingual Dataset for Corporate Promise Verification
di: Seki, Yohei, et al.
Pubblicazione: (2024)
di: Seki, Yohei, et al.
Pubblicazione: (2024)
AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
di: Hu, Yuxuan, et al.
Pubblicazione: (2026)
di: Hu, Yuxuan, et al.
Pubblicazione: (2026)
Learning Efficient Guardrails for Compliance
di: Wen, Xiaofei, et al.
Pubblicazione: (2025)
di: Wen, Xiaofei, et al.
Pubblicazione: (2025)
Mixup Model Merge: Enhancing Model Merging Performance through Randomized Linear Interpolation
di: Zhou, Yue, et al.
Pubblicazione: (2025)
di: Zhou, Yue, et al.
Pubblicazione: (2025)
TrafficRAG: A Multimodal RAG Framework for Traffic Accident Liability Determination
di: Li, Xu, et al.
Pubblicazione: (2026)
di: Li, Xu, et al.
Pubblicazione: (2026)
Dodo: Dynamic Contextual Compression for Decoder-only LMs
di: Qin, Guanghui, et al.
Pubblicazione: (2023)
di: Qin, Guanghui, et al.
Pubblicazione: (2023)
ConPET: Continual Parameter-Efficient Tuning for Large Language Models
di: Song, Chenyang, et al.
Pubblicazione: (2023)
di: Song, Chenyang, et al.
Pubblicazione: (2023)
A Knowledge Enhanced Learning and Semantic Composition Model for Multi-Claim Fact Checking
di: Wang, Shuai, et al.
Pubblicazione: (2021)
di: Wang, Shuai, et al.
Pubblicazione: (2021)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
di: Mi, Zhendong, et al.
Pubblicazione: (2025)
di: Mi, Zhendong, et al.
Pubblicazione: (2025)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
di: Liu, Fangxin, et al.
Pubblicazione: (2025)
di: Liu, Fangxin, et al.
Pubblicazione: (2025)
IWLV-Ramayana: A Sarga-Aligned Parallel Corpus of Valmiki's Ramayana Across Indian Languages
di: VP, Sumesh
Pubblicazione: (2026)
di: VP, Sumesh
Pubblicazione: (2026)
The Perplexity Paradox: Why Code Compresses Better Than Math in LLM Prompts
di: Johnson, Warren
Pubblicazione: (2026)
di: Johnson, Warren
Pubblicazione: (2026)
Xinyu: An Efficient LLM-based System for Commentary Generation
di: Wu, Yiquan, et al.
Pubblicazione: (2024)
di: Wu, Yiquan, et al.
Pubblicazione: (2024)
Large Language Model (LLM) Bias Index -- LLMBI
di: Oketunji, Abiodun Finbarrs, et al.
Pubblicazione: (2023)
di: Oketunji, Abiodun Finbarrs, et al.
Pubblicazione: (2023)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
di: Saji, Alan, et al.
Pubblicazione: (2025)
di: Saji, Alan, et al.
Pubblicazione: (2025)
Pre-trained Language Model with Prompts for Temporal Knowledge Graph Completion
di: Xu, Wenjie, et al.
Pubblicazione: (2023)
di: Xu, Wenjie, et al.
Pubblicazione: (2023)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
di: Palit, Sayon, et al.
Pubblicazione: (2025)
di: Palit, Sayon, et al.
Pubblicazione: (2025)
Homogeneous Keys, Heterogeneous Values: Exploiting Local KV Cache Asymmetry for Long-Context LLMs
di: Cui, Wanyun, et al.
Pubblicazione: (2025)
di: Cui, Wanyun, et al.
Pubblicazione: (2025)
Character-Level Transformer for Tajik-Persian Transliteration with a Parallel Lexical Corpus
di: Arabov, Mullosharaf K.
Pubblicazione: (2026)
di: Arabov, Mullosharaf K.
Pubblicazione: (2026)
An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs
di: Zhu, Qian, et al.
Pubblicazione: (2026)
di: Zhu, Qian, et al.
Pubblicazione: (2026)
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
di: Karinshak, Elise, et al.
Pubblicazione: (2024)
di: Karinshak, Elise, et al.
Pubblicazione: (2024)
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
di: Wang, Xintao, et al.
Pubblicazione: (2026)
di: Wang, Xintao, et al.
Pubblicazione: (2026)
Heimdall: test-time scaling on the generative verification
di: Shi, Wenlei, et al.
Pubblicazione: (2025)
di: Shi, Wenlei, et al.
Pubblicazione: (2025)
Executing Arithmetic: Fine-Tuning Large Language Models as Turing Machines
di: Lai, Junyu, et al.
Pubblicazione: (2024)
di: Lai, Junyu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search for LLM Inference
di: Dong, Wenxin, et al.
Pubblicazione: (2026) -
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025) -
SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning
di: Chang, Edward Y., et al.
Pubblicazione: (2025) -
Active Context Compression: Autonomous Memory Management in LLM Agents
di: Verma, Nikhil
Pubblicazione: (2026) -
Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment
di: Chang, Edward Y.
Pubblicazione: (2026)