Efficient LLM-based Advertising via Model Compression and Parallel Verification
Fuente:
arXiv
Guardado en:
| Autores principales: | Dong, Wenxin, Gao, Chang, Yu, Guanghui, Jiao, Xuewu, Hu, Mingqing, Fu, Qiang, Xu, Peng, Wei, Penghui, Xu, Hui, Xing, Yue, Li, Shuanglong, Liu, Lin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search for LLM Inference
por: Dong, Wenxin, et al.
Publicado: (2026)
por: Dong, Wenxin, et al.
Publicado: (2026)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
por: Peters, Sydney, et al.
Publicado: (2025)
por: Peters, Sydney, et al.
Publicado: (2025)
SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning
por: Chang, Edward Y., et al.
Publicado: (2025)
por: Chang, Edward Y., et al.
Publicado: (2025)
Active Context Compression: Autonomous Memory Management in LLM Agents
por: Verma, Nikhil
Publicado: (2026)
por: Verma, Nikhil
Publicado: (2026)
Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment
por: Chang, Edward Y.
Publicado: (2026)
por: Chang, Edward Y.
Publicado: (2026)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
por: Souza, Débora, et al.
Publicado: (2026)
por: Souza, Débora, et al.
Publicado: (2026)
LLM-Based SQL Generation: Prompting, Self-Refinement, and Adaptive Weighted Majority Voting
por: Yang, Yu-Jie, et al.
Publicado: (2026)
por: Yang, Yu-Jie, et al.
Publicado: (2026)
On the Challenges of Creating Datasets for Analyzing Commercial Sex Advertisements to Assess Human Trafficking Risk and Organized Activity
por: Rivas, Pablo, et al.
Publicado: (2024)
por: Rivas, Pablo, et al.
Publicado: (2024)
EVINCE: Optimizing Multi-LLM Dialogues Using Conditional Statistics and Information Theory
por: Chang, Edward Y.
Publicado: (2024)
por: Chang, Edward Y.
Publicado: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
por: Ashuach, Tomer, et al.
Publicado: (2025)
por: Ashuach, Tomer, et al.
Publicado: (2025)
LLM-based Automated Theorem Proving Hinges on Scalable Synthetic Data Generation
por: Lai, Junyu, et al.
Publicado: (2025)
por: Lai, Junyu, et al.
Publicado: (2025)
ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning
por: Chang, Edward Y., et al.
Publicado: (2025)
por: Chang, Edward Y., et al.
Publicado: (2025)
From Classification to Ranking: Enhancing LLM Reasoning Capabilities for MBTI Personality Detection
por: Cao, Yuan, et al.
Publicado: (2026)
por: Cao, Yuan, et al.
Publicado: (2026)
Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
por: Johnson, Warren
Publicado: (2026)
por: Johnson, Warren
Publicado: (2026)
PLM: Efficient Peripheral Language Models Hardware-Co-Designed for Ubiquitous Computing
por: Deng, Cheng, et al.
Publicado: (2025)
por: Deng, Cheng, et al.
Publicado: (2025)
LLM Vocabulary Compression for Low-Compute Environments
por: Vennam, Sreeram, et al.
Publicado: (2024)
por: Vennam, Sreeram, et al.
Publicado: (2024)
ML-Promise: A Multilingual Dataset for Corporate Promise Verification
por: Seki, Yohei, et al.
Publicado: (2024)
por: Seki, Yohei, et al.
Publicado: (2024)
AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
por: Hu, Yuxuan, et al.
Publicado: (2026)
por: Hu, Yuxuan, et al.
Publicado: (2026)
Learning Efficient Guardrails for Compliance
por: Wen, Xiaofei, et al.
Publicado: (2025)
por: Wen, Xiaofei, et al.
Publicado: (2025)
Mixup Model Merge: Enhancing Model Merging Performance through Randomized Linear Interpolation
por: Zhou, Yue, et al.
Publicado: (2025)
por: Zhou, Yue, et al.
Publicado: (2025)
TrafficRAG: A Multimodal RAG Framework for Traffic Accident Liability Determination
por: Li, Xu, et al.
Publicado: (2026)
por: Li, Xu, et al.
Publicado: (2026)
Dodo: Dynamic Contextual Compression for Decoder-only LMs
por: Qin, Guanghui, et al.
Publicado: (2023)
por: Qin, Guanghui, et al.
Publicado: (2023)
ConPET: Continual Parameter-Efficient Tuning for Large Language Models
por: Song, Chenyang, et al.
Publicado: (2023)
por: Song, Chenyang, et al.
Publicado: (2023)
A Knowledge Enhanced Learning and Semantic Composition Model for Multi-Claim Fact Checking
por: Wang, Shuai, et al.
Publicado: (2021)
por: Wang, Shuai, et al.
Publicado: (2021)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
por: Mi, Zhendong, et al.
Publicado: (2025)
por: Mi, Zhendong, et al.
Publicado: (2025)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
por: Liu, Fangxin, et al.
Publicado: (2025)
por: Liu, Fangxin, et al.
Publicado: (2025)
IWLV-Ramayana: A Sarga-Aligned Parallel Corpus of Valmiki's Ramayana Across Indian Languages
por: VP, Sumesh
Publicado: (2026)
por: VP, Sumesh
Publicado: (2026)
The Perplexity Paradox: Why Code Compresses Better Than Math in LLM Prompts
por: Johnson, Warren
Publicado: (2026)
por: Johnson, Warren
Publicado: (2026)
Xinyu: An Efficient LLM-based System for Commentary Generation
por: Wu, Yiquan, et al.
Publicado: (2024)
por: Wu, Yiquan, et al.
Publicado: (2024)
Large Language Model (LLM) Bias Index -- LLMBI
por: Oketunji, Abiodun Finbarrs, et al.
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs, et al.
Publicado: (2023)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
por: Saji, Alan, et al.
Publicado: (2025)
por: Saji, Alan, et al.
Publicado: (2025)
Pre-trained Language Model with Prompts for Temporal Knowledge Graph Completion
por: Xu, Wenjie, et al.
Publicado: (2023)
por: Xu, Wenjie, et al.
Publicado: (2023)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
por: Palit, Sayon, et al.
Publicado: (2025)
por: Palit, Sayon, et al.
Publicado: (2025)
Homogeneous Keys, Heterogeneous Values: Exploiting Local KV Cache Asymmetry for Long-Context LLMs
por: Cui, Wanyun, et al.
Publicado: (2025)
por: Cui, Wanyun, et al.
Publicado: (2025)
Character-Level Transformer for Tajik-Persian Transliteration with a Parallel Lexical Corpus
por: Arabov, Mullosharaf K.
Publicado: (2026)
por: Arabov, Mullosharaf K.
Publicado: (2026)
An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs
por: Zhu, Qian, et al.
Publicado: (2026)
por: Zhu, Qian, et al.
Publicado: (2026)
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
por: Karinshak, Elise, et al.
Publicado: (2024)
por: Karinshak, Elise, et al.
Publicado: (2024)
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
por: Wang, Xintao, et al.
Publicado: (2026)
por: Wang, Xintao, et al.
Publicado: (2026)
Heimdall: test-time scaling on the generative verification
por: Shi, Wenlei, et al.
Publicado: (2025)
por: Shi, Wenlei, et al.
Publicado: (2025)
Executing Arithmetic: Fine-Tuning Large Language Models as Turing Machines
por: Lai, Junyu, et al.
Publicado: (2024)
por: Lai, Junyu, et al.
Publicado: (2024)
Ejemplares similares
-
Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search for LLM Inference
por: Dong, Wenxin, et al.
Publicado: (2026) -
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
por: Peters, Sydney, et al.
Publicado: (2025) -
SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning
por: Chang, Edward Y., et al.
Publicado: (2025) -
Active Context Compression: Autonomous Memory Management in LLM Agents
por: Verma, Nikhil
Publicado: (2026) -
Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment
por: Chang, Edward Y.
Publicado: (2026)