Is (Selective) Round-To-Nearest Quantization All You Need?
Fuente:
arXiv
Salvato in:
| Autore principale: | Kogan, Alex |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
KV Cache Optimization Strategies for Scalable and Efficient LLM Inference
di: Xu, Yichun, et al.
Pubblicazione: (2026)
di: Xu, Yichun, et al.
Pubblicazione: (2026)
GALA: Multimodal Graph Alignment for Bug Localization in Automated Program Repair
di: Liu, Zhuoyao, et al.
Pubblicazione: (2026)
di: Liu, Zhuoyao, et al.
Pubblicazione: (2026)
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
di: Dua, Karan, et al.
Pubblicazione: (2025)
di: Dua, Karan, et al.
Pubblicazione: (2025)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
di: Li, Xiangchen, et al.
Pubblicazione: (2026)
di: Li, Xiangchen, et al.
Pubblicazione: (2026)
METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark
di: Yang, Xu, et al.
Pubblicazione: (2025)
di: Yang, Xu, et al.
Pubblicazione: (2025)
Unpacking the Eye of the Beholder: Social Location, Identity, and the Moving Target of Political Perspectives
di: Sirotkina, Elena
Pubblicazione: (2026)
di: Sirotkina, Elena
Pubblicazione: (2026)
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
di: Li, Xiangchen, et al.
Pubblicazione: (2026)
di: Li, Xiangchen, et al.
Pubblicazione: (2026)
Neural Attention: A Novel Mechanism for Enhanced Expressive Power in Transformer Models
di: DiGiugno, Andrew, et al.
Pubblicazione: (2025)
di: DiGiugno, Andrew, et al.
Pubblicazione: (2025)
Leum-VL Technical Report
di: He, Yuxuan, et al.
Pubblicazione: (2026)
di: He, Yuxuan, et al.
Pubblicazione: (2026)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
di: Liu, Bingnan, et al.
Pubblicazione: (2026)
di: Liu, Bingnan, et al.
Pubblicazione: (2026)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
di: Chen, Yuangong, et al.
Pubblicazione: (2026)
di: Chen, Yuangong, et al.
Pubblicazione: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
Fidelity Probes for Specification--Code Alignment
di: Erata, Ferhat, et al.
Pubblicazione: (2026)
di: Erata, Ferhat, et al.
Pubblicazione: (2026)
GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models
di: Hacheme, Gilles Quentin, et al.
Pubblicazione: (2025)
di: Hacheme, Gilles Quentin, et al.
Pubblicazione: (2025)
A Genealogy of Foundation Models in Remote Sensing
di: Lane, Kevin, et al.
Pubblicazione: (2025)
di: Lane, Kevin, et al.
Pubblicazione: (2025)
Do All Vision Transformers Need Registers? A Cross-Architectural Reassessment
di: Baxevanakis, Spiros, et al.
Pubblicazione: (2026)
di: Baxevanakis, Spiros, et al.
Pubblicazione: (2026)
StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
di: Oliveira, Daniel, et al.
Pubblicazione: (2026)
di: Oliveira, Daniel, et al.
Pubblicazione: (2026)
Supervised Embedded Methods for Hyperspectral Band Selection
di: Zimmer, Yaniv, et al.
Pubblicazione: (2024)
di: Zimmer, Yaniv, et al.
Pubblicazione: (2024)
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
di: Shahin, Nada, et al.
Pubblicazione: (2025)
di: Shahin, Nada, et al.
Pubblicazione: (2025)
Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs
di: Sada, Mohammad Firas, et al.
Pubblicazione: (2025)
di: Sada, Mohammad Firas, et al.
Pubblicazione: (2025)
Transfer-learning for video classification: Video Swin Transformer on multiple domains
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2022)
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2022)
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
di: Nemitz, Jonathan, et al.
Pubblicazione: (2026)
di: Nemitz, Jonathan, et al.
Pubblicazione: (2026)
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
di: Hou, Zhiyi, et al.
Pubblicazione: (2025)
di: Hou, Zhiyi, et al.
Pubblicazione: (2025)
Making a Pipeline Production-Ready: Challenges and Lessons Learned in the Healthcare Domain
di: Lawand, Daniel Angelo Esteves, et al.
Pubblicazione: (2025)
di: Lawand, Daniel Angelo Esteves, et al.
Pubblicazione: (2025)
SPIRA: Building an Intelligent System for Respiratory Insufficiency Detection
di: Ferreira, Renato Cordeiro, et al.
Pubblicazione: (2025)
di: Ferreira, Renato Cordeiro, et al.
Pubblicazione: (2025)
Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference
di: Gao, Yifei, et al.
Pubblicazione: (2026)
di: Gao, Yifei, et al.
Pubblicazione: (2026)
Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models
di: Wang, Junxin, et al.
Pubblicazione: (2026)
di: Wang, Junxin, et al.
Pubblicazione: (2026)
Using Deep Learning to Generate Semantically Correct Hindi Captions
di: Khan, Wasim Akram, et al.
Pubblicazione: (2026)
di: Khan, Wasim Akram, et al.
Pubblicazione: (2026)
RAPTOR-AI for Disaster OODA Loop: Hierarchical Multimodal RAG with Experience-Driven Agentic Decision-Making
di: Yasuno, Takato
Pubblicazione: (2026)
di: Yasuno, Takato
Pubblicazione: (2026)
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
di: Shahin, Nada, et al.
Pubblicazione: (2025)
di: Shahin, Nada, et al.
Pubblicazione: (2025)
Learning, Fast and Slow: Towards LLMs That Adapt Continually
di: Tiwari, Rishabh, et al.
Pubblicazione: (2026)
di: Tiwari, Rishabh, et al.
Pubblicazione: (2026)
Knowledge Graphs as the Missing Data Layer for LLM-Based Industrial Asset Operations
di: Mandarapu, Madhulatha, et al.
Pubblicazione: (2026)
di: Mandarapu, Madhulatha, et al.
Pubblicazione: (2026)
DP-2Stage: Adapting Language Models as Differentially Private Tabular Data Generators
di: Afonja, Tejumade, et al.
Pubblicazione: (2024)
di: Afonja, Tejumade, et al.
Pubblicazione: (2024)
LLMs as Idiomatic Decompilers: Recovering High-Level Code from x86-64 Assembly for Dart
di: Abualazm, Raafat, et al.
Pubblicazione: (2026)
di: Abualazm, Raafat, et al.
Pubblicazione: (2026)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
di: Othman, Refat
Pubblicazione: (2026)
di: Othman, Refat
Pubblicazione: (2026)
Predicting the descent into extremism and terrorism
di: Lane, R. O., et al.
Pubblicazione: (2025)
di: Lane, R. O., et al.
Pubblicazione: (2025)
Estimating optical vegetation indices and biophysical variables for temperate forests with Sentinel-1 SAR data using machine learning techniques: A case study for Czechia
di: Paluba, Daniel, et al.
Pubblicazione: (2023)
di: Paluba, Daniel, et al.
Pubblicazione: (2023)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
Context-Instrumental Data Distillation for Kubernetes Manifest Generation: Method and Experimental Evaluation
di: Kozachok, Andrey, et al.
Pubblicazione: (2026)
di: Kozachok, Andrey, et al.
Pubblicazione: (2026)
VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations
di: Gautam, Sushant, et al.
Pubblicazione: (2026)
di: Gautam, Sushant, et al.
Pubblicazione: (2026)
Documenti analoghi
-
KV Cache Optimization Strategies for Scalable and Efficient LLM Inference
di: Xu, Yichun, et al.
Pubblicazione: (2026) -
GALA: Multimodal Graph Alignment for Bug Localization in Automated Program Repair
di: Liu, Zhuoyao, et al.
Pubblicazione: (2026) -
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
di: Dua, Karan, et al.
Pubblicazione: (2025) -
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
di: Li, Xiangchen, et al.
Pubblicazione: (2026) -
METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark
di: Yang, Xu, et al.
Pubblicazione: (2025)