WebLLM: A High-Performance In-Browser LLM Inference Engine
Fuente:
arXiv
Guardado en:
| Autores principales: | Ruan, Charlie F., Qin, Yucheng, Parthasarathy, Akaash R., Zhou, Xun, Lai, Ruihang, Jin, Hongyi, Dong, Yixin, Hou, Bohan, Yu, Meng-Shiun, Zhai, Yiyan, Agarwal, Sudeep, Cao, Hangrui, Feng, Siyuan, Chen, Tianqi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models
por: Dong, Yixin, et al.
Publicado: (2024)
por: Dong, Yixin, et al.
Publicado: (2024)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
por: Ye, Zihao, et al.
Publicado: (2025)
por: Ye, Zihao, et al.
Publicado: (2025)
Productively Deploying Emerging Models on Emerging Platforms: A Top-Down Approach for Testing and Debugging
por: Feng, Siyuan, et al.
Publicado: (2024)
por: Feng, Siyuan, et al.
Publicado: (2024)
PithTrain: A Compact and Agent-Native MoE Training System
por: Lai, Ruihang, et al.
Publicado: (2026)
por: Lai, Ruihang, et al.
Publicado: (2026)
FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems
por: Xing, Shanli, et al.
Publicado: (2026)
por: Xing, Shanli, et al.
Publicado: (2026)
A System for Microserving of LLMs
por: Jin, Hongyi, et al.
Publicado: (2024)
por: Jin, Hongyi, et al.
Publicado: (2024)
Anatomizing Deep Learning Inference in Web Browsers
por: Wang, Qipeng, et al.
Publicado: (2024)
por: Wang, Qipeng, et al.
Publicado: (2024)
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
por: Anupam, Sagnik, et al.
Publicado: (2025)
por: Anupam, Sagnik, et al.
Publicado: (2025)
Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers
por: Maczan, Jędrzej
Publicado: (2026)
por: Maczan, Jędrzej
Publicado: (2026)
Local deployment of large-scale music AI models on commodity hardware
por: Zhou, Xun, et al.
Publicado: (2024)
por: Zhou, Xun, et al.
Publicado: (2024)
Axe: A Simple Unified Layout Abstraction for Machine Learning Compilers
por: Hou, Bohan, et al.
Publicado: (2026)
por: Hou, Bohan, et al.
Publicado: (2026)
Biotic Browser: Applying StreamingLLM as a Persistent Web Browsing Co-Pilot
por: Dunnell, Kevin F., et al.
Publicado: (2024)
por: Dunnell, Kevin F., et al.
Publicado: (2024)
Client-Side Zero-Shot LLM Inference for Comprehensive In-Browser URL Analysis
por: Cohen, Avihay
Publicado: (2025)
por: Cohen, Avihay
Publicado: (2025)
RTP-LLM: High-Performance Alibaba LLM Inference Engine
por: Tan, Boyu, et al.
Publicado: (2026)
por: Tan, Boyu, et al.
Publicado: (2026)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
por: Yu, Bohan, et al.
Publicado: (2025)
por: Yu, Bohan, et al.
Publicado: (2025)
A First Look at Bugs in LLM Inference Engines
por: Liu, Mugeng, et al.
Publicado: (2025)
por: Liu, Mugeng, et al.
Publicado: (2025)
In-Browser LLM-Guided Fuzzing for Real-Time Prompt Injection Testing in Agentic AI Browsers
por: Cohen, Avihay
Publicado: (2025)
por: Cohen, Avihay
Publicado: (2025)
Causal Inference for Human-Language Model Collaboration
por: Zhang, Bohan, et al.
Publicado: (2024)
por: Zhang, Bohan, et al.
Publicado: (2024)
Browser Fingerprinting Using WebAssembly
por: Guri, Mordechai, et al.
Publicado: (2025)
por: Guri, Mordechai, et al.
Publicado: (2025)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
por: Cheng, Ke, et al.
Publicado: (2024)
por: Cheng, Ke, et al.
Publicado: (2024)
ExaCraft: Dynamic Learning Context Adaptation for Personalized Educational Examples
por: Chatterjee, Akaash, et al.
Publicado: (2025)
por: Chatterjee, Akaash, et al.
Publicado: (2025)
TIDE: Temporal Incremental Draft Engine for Self-Improving LLM Inference
por: Park, Jiyoung, et al.
Publicado: (2026)
por: Park, Jiyoung, et al.
Publicado: (2026)
Orion: Characterizing and Programming Apple's Neural Engine for LLM Training and Inference
por: Kumaresan, Ramchand
Publicado: (2026)
por: Kumaresan, Ramchand
Publicado: (2026)
The BrowserGym Ecosystem for Web Agent Research
por: De Chezelles, Thibault Le Sellier, et al.
Publicado: (2024)
por: De Chezelles, Thibault Le Sellier, et al.
Publicado: (2024)
WAAA! Web Adversaries Against Agentic Browsers
por: Datta, Sohom, et al.
Publicado: (2026)
por: Datta, Sohom, et al.
Publicado: (2026)
User Profiles: The Achilles' Heel of Web Browsers
por: Somé, Dolière Francis, et al.
Publicado: (2025)
por: Somé, Dolière Francis, et al.
Publicado: (2025)
Blackbox Dataset Inference for LLM
por: Zhou, Ruikai, et al.
Publicado: (2025)
por: Zhou, Ruikai, et al.
Publicado: (2025)
Research on Multi-hop Inference Optimization of LLM Based on MQUAKE Framework
por: Liang, Zucheng, et al.
Publicado: (2025)
por: Liang, Zucheng, et al.
Publicado: (2025)
Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces
por: Lugoloobi, William, et al.
Publicado: (2026)
por: Lugoloobi, William, et al.
Publicado: (2026)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
por: He, Yongchao, et al.
Publicado: (2025)
por: He, Yongchao, et al.
Publicado: (2025)
SearchRAG: Can Search Engines Be Helpful for LLM-based Medical Question Answering?
por: Shi, Yucheng, et al.
Publicado: (2025)
por: Shi, Yucheng, et al.
Publicado: (2025)
SparQ Attention: Bandwidth-Efficient LLM Inference
por: Ribar, Luka, et al.
Publicado: (2023)
por: Ribar, Luka, et al.
Publicado: (2023)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
por: Agarwal, Saurabh, et al.
Publicado: (2024)
por: Agarwal, Saurabh, et al.
Publicado: (2024)
XGrammar-2: Efficient Dynamic Structured Generation Engine for Agentic LLMs
por: Li, Linzhang, et al.
Publicado: (2026)
por: Li, Linzhang, et al.
Publicado: (2026)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
por: Zhao, Bohan, et al.
Publicado: (2025)
por: Zhao, Bohan, et al.
Publicado: (2025)
How Web Browsers Shape Users' Understanding of Networks.
por: Sheeran, Louise, et al.
Publicado: (2002)
por: Sheeran, Louise, et al.
Publicado: (2002)
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
por: Fu, Tianyu, et al.
Publicado: (2024)
por: Fu, Tianyu, et al.
Publicado: (2024)
WebANNS: Fast and Efficient Approximate Nearest Neighbor Search in Web Browsers
por: Liu, Mugeng, et al.
Publicado: (2025)
por: Liu, Mugeng, et al.
Publicado: (2025)
BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
por: Yu, Tao, et al.
Publicado: (2025)
por: Yu, Tao, et al.
Publicado: (2025)
Verify as You Go: An LLM-Powered Browser Extension for Fake News Detection
por: Sallami, Dorsaf, et al.
Publicado: (2026)
por: Sallami, Dorsaf, et al.
Publicado: (2026)
Ejemplares similares
-
XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models
por: Dong, Yixin, et al.
Publicado: (2024) -
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
por: Ye, Zihao, et al.
Publicado: (2025) -
Productively Deploying Emerging Models on Emerging Platforms: A Top-Down Approach for Testing and Debugging
por: Feng, Siyuan, et al.
Publicado: (2024) -
PithTrain: A Compact and Agent-Native MoE Training System
por: Lai, Ruihang, et al.
Publicado: (2026) -
FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems
por: Xing, Shanli, et al.
Publicado: (2026)