Partial Convolution Meets Visual Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Haiduo, Yang, Fuwei, Li, Dong, Liu, Ji, Tian, Lu, Peng, Jinzhang, Ren, Pengju, Barsoum, Emad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SpecVLM: Fast Speculative Decoding in Vision-Language Models
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
Partial Channel Network: Compute Fewer, Perform Better
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
Fast Occupancy Network
von: Lu, Mingjie, et al.
Veröffentlicht: (2024)
von: Lu, Mingjie, et al.
Veröffentlicht: (2024)
Nearly Lossless Adaptive Bit Switching
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
GeGS-PCR: Effective and Robust 3D Point Cloud Registration with Two-Stage Color-Enhanced Geometric-3DGS Fusion
von: Tian, Jiayi, et al.
Veröffentlicht: (2026)
von: Tian, Jiayi, et al.
Veröffentlicht: (2026)
KernelDNA: Dynamic Kernel Sharing via Decoupled Naive Adapters
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
DeepKD: A Deeply Decoupled and Denoised Knowledge Distillation Trainer
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
ICAS: IP Adapter and ControlNet-based Attention Structure for Multi-Subject Style Transfer Optimization
von: Liu, Fuwei
Veröffentlicht: (2025)
von: Liu, Fuwei
Veröffentlicht: (2025)
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
von: Ke, Wenjin, et al.
Veröffentlicht: (2025)
von: Ke, Wenjin, et al.
Veröffentlicht: (2025)
Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion
von: Singh, Shivam, et al.
Veröffentlicht: (2026)
von: Singh, Shivam, et al.
Veröffentlicht: (2026)
UPDP: A Unified Progressive Depth Pruner for CNN and Vision Transformer
von: Liu, Ji, et al.
Veröffentlicht: (2024)
von: Liu, Ji, et al.
Veröffentlicht: (2024)
Sparse Laneformer
von: Liu, Ji, et al.
Veröffentlicht: (2024)
von: Liu, Ji, et al.
Veröffentlicht: (2024)
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2026)
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2026)
DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking
von: Haridas, Akash, et al.
Veröffentlicht: (2026)
von: Haridas, Akash, et al.
Veröffentlicht: (2026)
Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
EGSRAL: An Enhanced 3D Gaussian Splatting based Renderer with Automated Labeling for Large-Scale Driving Scene
von: Huo, Yixiong, et al.
Veröffentlicht: (2024)
von: Huo, Yixiong, et al.
Veröffentlicht: (2024)
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
CLAReSNet: When Convolution Meets Latent Attention for Hyperspectral Image Classification
von: Bandyopadhyay, Asmit, et al.
Veröffentlicht: (2025)
von: Bandyopadhyay, Asmit, et al.
Veröffentlicht: (2025)
Learning from Online Videos at Inference Time for Computer-Use Agents
von: Liu, Yujian, et al.
Veröffentlicht: (2025)
von: Liu, Yujian, et al.
Veröffentlicht: (2025)
VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping
von: Dong, Haotian, et al.
Veröffentlicht: (2025)
von: Dong, Haotian, et al.
Veröffentlicht: (2025)
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
von: Zheng, Zirui, et al.
Veröffentlicht: (2025)
von: Zheng, Zirui, et al.
Veröffentlicht: (2025)
Revisiting the Integration of Convolution and Attention for Vision Backbone
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
Variational Partial Group Convolutions for Input-Aware Partial Equivariance of Rotations and Color-Shifts
von: Kim, Hyunsu, et al.
Veröffentlicht: (2024)
von: Kim, Hyunsu, et al.
Veröffentlicht: (2024)
LADDER: An Efficient Framework for Video Frame Interpolation
von: Shen, Tong, et al.
Veröffentlicht: (2024)
von: Shen, Tong, et al.
Veröffentlicht: (2024)
MonoGS++: Fast and Accurate Monocular RGB Gaussian SLAM
von: Li, Renwu, et al.
Veröffentlicht: (2025)
von: Li, Renwu, et al.
Veröffentlicht: (2025)
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
von: Lin, Jingyang, et al.
Veröffentlicht: (2026)
von: Lin, Jingyang, et al.
Veröffentlicht: (2026)
SwinECAT: A Transformer-based fundus disease classification model with Shifted Window Attention and Efficient Channel Attention
von: Gu, Peiran, et al.
Veröffentlicht: (2025)
von: Gu, Peiran, et al.
Veröffentlicht: (2025)
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
von: Zhang, Qizhe, et al.
Veröffentlicht: (2024)
von: Zhang, Qizhe, et al.
Veröffentlicht: (2024)
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
Bladder Vessel Segmentation using a Hybrid Attention-Convolution Framework
von: Krauß, Franziska, et al.
Veröffentlicht: (2026)
von: Krauß, Franziska, et al.
Veröffentlicht: (2026)
TSGCNeXt: Dynamic-Static Multi-Graph Convolution for Efficient Skeleton-Based Action Recognition with Long-term Learning Potential
von: Liu, Dongjingdin, et al.
Veröffentlicht: (2023)
von: Liu, Dongjingdin, et al.
Veröffentlicht: (2023)
Object Isolated Attention for Consistent Story Visualization
von: Luo, Xiangyang, et al.
Veröffentlicht: (2025)
von: Luo, Xiangyang, et al.
Veröffentlicht: (2025)
VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning
von: Dong, Mingkang, et al.
Veröffentlicht: (2026)
von: Dong, Mingkang, et al.
Veröffentlicht: (2026)
Attention at Rest Stays at Rest: Breaking Visual Inertia for Cognitive Hallucination Mitigation
von: Gong, Boyang, et al.
Veröffentlicht: (2026)
von: Gong, Boyang, et al.
Veröffentlicht: (2026)
Enhancing Satellite Object Localization with Dilated Convolutions and Attention-aided Spatial Pooling
von: Mostafa, Seraj Al Mahmud, et al.
Veröffentlicht: (2025)
von: Mostafa, Seraj Al Mahmud, et al.
Veröffentlicht: (2025)
ECViT: Efficient Convolutional Vision Transformer with Local-Attention and Multi-scale Stages
von: Qian, Zhoujie
Veröffentlicht: (2025)
von: Qian, Zhoujie
Veröffentlicht: (2025)
SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
RELO: Reinforcement Learning to Localize for Visual Object Tracking
von: Chen, Xin, et al.
Veröffentlicht: (2026)
von: Chen, Xin, et al.
Veröffentlicht: (2026)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
von: Ge, Yuyao, et al.
Veröffentlicht: (2025)
von: Ge, Yuyao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SpecVLM: Fast Speculative Decoding in Vision-Language Models
von: Huang, Haiduo, et al.
Veröffentlicht: (2025) -
Partial Channel Network: Compute Fewer, Perform Better
von: Huang, Haiduo, et al.
Veröffentlicht: (2025) -
Fast Occupancy Network
von: Lu, Mingjie, et al.
Veröffentlicht: (2024) -
Nearly Lossless Adaptive Bit Switching
von: Huang, Haiduo, et al.
Veröffentlicht: (2025) -
GeGS-PCR: Effective and Robust 3D Point Cloud Registration with Two-Stage Color-Enhanced Geometric-3DGS Fusion
von: Tian, Jiayi, et al.
Veröffentlicht: (2026)