Saved in:
| Main Authors: | Qin, Haotong, Hu, Cheng, Magno, Michele |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.07627 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PicoSAM2: Low-Latency Segmentation In-Sensor for Edge Vision Applications
by: Bonazzi, Pietro, et al.
Published: (2025)
by: Bonazzi, Pietro, et al.
Published: (2025)
Event-Based Vision in Space: Applications, Trends, and Future Directions
by: Capogrosso, Luigi, et al.
Published: (2026)
by: Capogrosso, Luigi, et al.
Published: (2026)
PicoSAM3: Real-Time In-Sensor Region-of-Interest Segmentation
by: Bonazzi, Pietro, et al.
Published: (2026)
by: Bonazzi, Pietro, et al.
Published: (2026)
RGB-Event Fusion with Self-Attention for Collision Prediction
by: Bonazzi, Pietro, et al.
Published: (2025)
by: Bonazzi, Pietro, et al.
Published: (2025)
BiVM: Accurate Binarized Neural Network for Efficient Video Matting
by: Qin, Haotong, et al.
Published: (2025)
by: Qin, Haotong, et al.
Published: (2025)
Q-SAM2: Accurate Quantization for Segment Anything Model 2
by: Farronato, Nicola, et al.
Published: (2025)
by: Farronato, Nicola, et al.
Published: (2025)
BinaryDM: Accurate Weight Binarization for Efficient Diffusion Models
by: Zheng, Xingyu, et al.
Published: (2024)
by: Zheng, Xingyu, et al.
Published: (2024)
Q-MambaIR: Accurate Quantized Mamba for Efficient Image Restoration
by: Chen, Yujie, et al.
Published: (2025)
by: Chen, Yujie, et al.
Published: (2025)
EVLM: An Efficient Vision-Language Model for Visual Understanding
by: Chen, Kaibing, et al.
Published: (2024)
by: Chen, Kaibing, et al.
Published: (2024)
QuantSR+: Pushing the Limit of Quantized Image Super-Resolution Networks
by: Qin, Haotong, et al.
Published: (2026)
by: Qin, Haotong, et al.
Published: (2026)
Post-Training Quantization for Video Matting
by: Zhu, Tianrui, et al.
Published: (2025)
by: Zhu, Tianrui, et al.
Published: (2025)
Quantized Visual Geometry Grounded Transformer
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
First-Order Error Matters: Accurate Compensation for Quantized Large Language Models
by: Zheng, Xingyu, et al.
Published: (2025)
by: Zheng, Xingyu, et al.
Published: (2025)
MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion Models
by: Feng, Weilun, et al.
Published: (2024)
by: Feng, Weilun, et al.
Published: (2024)
Temporal-Guided Visual Foundation Models for Event-Based Vision
by: Xia, Ruihao, et al.
Published: (2025)
by: Xia, Ruihao, et al.
Published: (2025)
Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models
by: Zhong, Yi, et al.
Published: (2026)
by: Zhong, Yi, et al.
Published: (2026)
Vision-Language Attribute Disentanglement and Reinforcement for Lifelong Person Re-Identification
by: Xu, Kunlun, et al.
Published: (2026)
by: Xu, Kunlun, et al.
Published: (2026)
QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
S$^2$Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
Planar Velocity Estimation for Fast-Moving Mobile Robots Using Event-Based Optical Flow
by: Boyle, Liam, et al.
Published: (2025)
by: Boyle, Liam, et al.
Published: (2025)
Efficient and Accurate Downfacing Visual Inertial Odometry
by: Kühne, Jonas, et al.
Published: (2025)
by: Kühne, Jonas, et al.
Published: (2025)
Image Fusion via Vision-Language Model
by: Zhao, Zixiang, et al.
Published: (2024)
by: Zhao, Zixiang, et al.
Published: (2024)
Do Vision-Language Models Really Understand Visual Language?
by: Hou, Yifan, et al.
Published: (2024)
by: Hou, Yifan, et al.
Published: (2024)
Do Vision-Language Models Understand Visual Persuasiveness?
by: Park, Gyuwon
Published: (2025)
by: Park, Gyuwon
Published: (2025)
Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
TinyML Enhances CubeSat Mission Capabilities
by: Capogrosso, Luigi, et al.
Published: (2026)
by: Capogrosso, Luigi, et al.
Published: (2026)
ScVLM: Enhancing Vision-Language Model for Safety-Critical Event Understanding
by: Shi, Liang, et al.
Published: (2024)
by: Shi, Liang, et al.
Published: (2024)
Exploiting In-Sensor Computing for Energy-Efficient Earth Observation
by: Capogrosso, Luigi, et al.
Published: (2026)
by: Capogrosso, Luigi, et al.
Published: (2026)
EventFlash: Towards Efficient MLLMs for Event-Based Vision
by: Liu, Shaoyu, et al.
Published: (2026)
by: Liu, Shaoyu, et al.
Published: (2026)
WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching
by: Feng, Weilun, et al.
Published: (2026)
by: Feng, Weilun, et al.
Published: (2026)
EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
by: Xu, Wenhao, et al.
Published: (2025)
by: Xu, Wenhao, et al.
Published: (2025)
MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models
by: Niu, Junbo, et al.
Published: (2025)
by: Niu, Junbo, et al.
Published: (2025)
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
by: Zhang, Yuan, et al.
Published: (2024)
by: Zhang, Yuan, et al.
Published: (2024)
MambaVF: State Space Model for Efficient Video Fusion
by: Zhao, Zixiang, et al.
Published: (2026)
by: Zhao, Zixiang, et al.
Published: (2026)
FTerViT: Fully Ternary Vision Transformer
by: Ruciński, Szymon, et al.
Published: (2026)
by: Ruciński, Szymon, et al.
Published: (2026)
PatchCue: Enhancing Vision-Language Model Reasoning with Patch-Based Visual Cues
by: Qi, Yukun, et al.
Published: (2026)
by: Qi, Yukun, et al.
Published: (2026)
VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
by: Liu, Qing'an, et al.
Published: (2026)
by: Liu, Qing'an, et al.
Published: (2026)
Low Latency Visual Inertial Odometry with On-Sensor Accelerated Optical Flow for Resource-Constrained UAVs
by: Kühne, Jonas, et al.
Published: (2024)
by: Kühne, Jonas, et al.
Published: (2024)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
by: Liu, Hanqing, et al.
Published: (2026)
by: Liu, Hanqing, et al.
Published: (2026)
Similar Items
-
PicoSAM2: Low-Latency Segmentation In-Sensor for Edge Vision Applications
by: Bonazzi, Pietro, et al.
Published: (2025) -
Event-Based Vision in Space: Applications, Trends, and Future Directions
by: Capogrosso, Luigi, et al.
Published: (2026) -
PicoSAM3: Real-Time In-Sensor Region-of-Interest Segmentation
by: Bonazzi, Pietro, et al.
Published: (2026) -
RGB-Event Fusion with Self-Attention for Collision Prediction
by: Bonazzi, Pietro, et al.
Published: (2025) -
BiVM: Accurate Binarized Neural Network for Efficient Video Matting
by: Qin, Haotong, et al.
Published: (2025)