PAT-VCM: Plug-and-Play Auxiliary Tokens for Video Coding for Machines
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Wei, Wang, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Iterative Reconstruction Method for Dental Cone-Beam Computed Tomography with a Truncated Field of View
by: Park, Hyoung Suk, et al.
Published: (2025)
by: Park, Hyoung Suk, et al.
Published: (2025)
Multimodal Image Matching based on Frequency-domain Information of Local Energy Response
by: Yang, Meng, et al.
Published: (2025)
by: Yang, Meng, et al.
Published: (2025)
Adaptive Multi-resolution Hash-Encoding Framework for INR-based Dental CBCT Reconstruction with Truncated FOV
by: Park, Hyoung Suk, et al.
Published: (2025)
by: Park, Hyoung Suk, et al.
Published: (2025)
Neural Image Compression Using Masked Sparse Visual Representation
by: Jiang, Wei, et al.
Published: (2023)
by: Jiang, Wei, et al.
Published: (2023)
An Effective Approach to Minimize Error in Midpoint Ellipse Drawing Algorithm
by: Idrisi, M. Javed, et al.
Published: (2021)
by: Idrisi, M. Javed, et al.
Published: (2021)
Sliced-Wasserstein Estimation with Spherical Harmonics as Control Variates
by: Leluc, Rémi, et al.
Published: (2024)
by: Leluc, Rémi, et al.
Published: (2024)
ML-ASPA: A Contemplation of Machine Learning-based Acoustic Signal Processing Analysis for Sounds, & Strains Emerging Technology
by: Ali, Ratul, et al.
Published: (2023)
by: Ali, Ratul, et al.
Published: (2023)
Z-Order Transformer for Feed-Forward Gaussian Splatting
by: Wang, Can, et al.
Published: (2026)
by: Wang, Can, et al.
Published: (2026)
MotionFollower: Editing Video Motion via Lightweight Score-Guided Diffusion
by: Tu, Shuyuan, et al.
Published: (2024)
by: Tu, Shuyuan, et al.
Published: (2024)
Depth Priors in Removal Neural Radiance Fields
by: Guo, Zhihao, et al.
Published: (2024)
by: Guo, Zhihao, et al.
Published: (2024)
When Less is Enough: Adaptive Token Reduction for Efficient Image Representation
by: Allakhverdov, Eduard, et al.
Published: (2025)
by: Allakhverdov, Eduard, et al.
Published: (2025)
Interactive Image Selection and Training for Brain Tumor Segmentation Network
by: Cerqueira, Matheus A., et al.
Published: (2024)
by: Cerqueira, Matheus A., et al.
Published: (2024)
MvBody: Multi-View-Based Hybrid Transformer Using Optical 3D Body Scan for Explainable Cesarean Section Prediction
by: Cheng, Ruting, et al.
Published: (2025)
by: Cheng, Ruting, et al.
Published: (2025)
Performance Decay in Deepfake Detection: The Limitations of Training on Outdated Data
by: Richings, Jack, et al.
Published: (2025)
by: Richings, Jack, et al.
Published: (2025)
How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study
by: Brusnicki, Roberto, et al.
Published: (2026)
by: Brusnicki, Roberto, et al.
Published: (2026)
A Formal Analysis of Algorithms for Matroids and Greedoids
by: Abdulaziz, Mohammad, et al.
Published: (2025)
by: Abdulaziz, Mohammad, et al.
Published: (2025)
EatGAN: An Edge-Attention Guided Generative Adversarial Network for Single Image Super-Resolution
by: Rao, Penghao, et al.
Published: (2025)
by: Rao, Penghao, et al.
Published: (2025)
TWIG: Two-Step Image Generation using Segmentation Masks in Diffusion Models
by: Rakib, Mazharul Islam, et al.
Published: (2025)
by: Rakib, Mazharul Islam, et al.
Published: (2025)
Isolated Sign Language Recognition with Segmentation and Pose Estimation
by: Perkins, Daniel, et al.
Published: (2025)
by: Perkins, Daniel, et al.
Published: (2025)
Motion Consistency Loss for Monocular Visual Odometry with Attention-Based Deep Learning
by: Françani, André O., et al.
Published: (2024)
by: Françani, André O., et al.
Published: (2024)
Less Detail, Better Answers: Degradation-Driven Prompting for VQA
by: Han, Haoxuan, et al.
Published: (2026)
by: Han, Haoxuan, et al.
Published: (2026)
TALON: Test-time Adaptive Learning for On-the-Fly Category Discovery
by: Wu, Yanan, et al.
Published: (2026)
by: Wu, Yanan, et al.
Published: (2026)
Attire-Based Anomaly Detection in Restricted Areas Using YOLOv8 for Enhanced CCTV Security
by: B, Abdul Aziz A., et al.
Published: (2024)
by: B, Abdul Aziz A., et al.
Published: (2024)
Transformer-Based Model for Monocular Visual Odometry: A Video Understanding Approach
by: Françani, André O., et al.
Published: (2023)
by: Françani, André O., et al.
Published: (2023)
Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think
by: Tian, Jie, et al.
Published: (2025)
by: Tian, Jie, et al.
Published: (2025)
On Graded Deformations of The universal Enveloping Algebra of a Color Lie Algebra
by: Petit, Toukaiddine
Published: (2025)
by: Petit, Toukaiddine
Published: (2025)
Gaussian-Constrained LeJEPA Representations for Unsupervised Scene Discovery and Pose Consistency
by: Mostafa, Mohsen
Published: (2026)
by: Mostafa, Mohsen
Published: (2026)
EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges
by: Jon, Hyo Jin, et al.
Published: (2026)
by: Jon, Hyo Jin, et al.
Published: (2026)
SealD-NeRF: Interactive Pixel-Level Editing for Dynamic Scenes by Neural Radiance Fields
by: Huang, Zhentao, et al.
Published: (2024)
by: Huang, Zhentao, et al.
Published: (2024)
An inpainting approach to manipulate asymmetry in pre-operative breast images
by: Montenegro, Helena, et al.
Published: (2025)
by: Montenegro, Helena, et al.
Published: (2025)
QCFace: Image Quality Control for boosting Face Representation & Recognition
by: Doan-Ngo, Duc-Phuong, et al.
Published: (2025)
by: Doan-Ngo, Duc-Phuong, et al.
Published: (2025)
Real-time Object and Event Detection Service through Computer Vision and Edge Computing
by: Mendes, Marcos, et al.
Published: (2025)
by: Mendes, Marcos, et al.
Published: (2025)
Not all Views are Created Equal: Analyzing Viewpoint Instabilities in Vision Foundation Models
by: Michalkiewicz, Mateusz, et al.
Published: (2024)
by: Michalkiewicz, Mateusz, et al.
Published: (2024)
Black-box Adversarial Attacks on CNN-based SLAM Algorithms
by: Gkeka, Maria Rafaela, et al.
Published: (2025)
by: Gkeka, Maria Rafaela, et al.
Published: (2025)
Knowledge-Guided Failure Prediction: Detecting When Object Detectors Miss Safety-Critical Objects
by: Zimmermann, Jakob Paul, et al.
Published: (2026)
by: Zimmermann, Jakob Paul, et al.
Published: (2026)
Sequence Transferability and Task Order Selection in Continual Learning
by: Nguyen, Thinh, et al.
Published: (2025)
by: Nguyen, Thinh, et al.
Published: (2025)
Building Brain Tumor Segmentation Networks with User-Assisted Filter Estimation and Selection
by: Cerqueira, Matheus A., et al.
Published: (2024)
by: Cerqueira, Matheus A., et al.
Published: (2024)
Bench2FreeAD: A Benchmark for Vision-based End-to-end Navigation in Unstructured Robotic Environments
by: Peng, Yuhang, et al.
Published: (2025)
by: Peng, Yuhang, et al.
Published: (2025)
Multimodal Integration Challenges in Emotionally Expressive Child Avatars for Training Applications
by: Salehi, Pegah, et al.
Published: (2025)
by: Salehi, Pegah, et al.
Published: (2025)
CylinderPlane: Nested Cylinder Representation for 3D-aware Image Generation
by: Jia, Ru, et al.
Published: (2025)
by: Jia, Ru, et al.
Published: (2025)
Similar Items
-
An Iterative Reconstruction Method for Dental Cone-Beam Computed Tomography with a Truncated Field of View
by: Park, Hyoung Suk, et al.
Published: (2025) -
Multimodal Image Matching based on Frequency-domain Information of Local Energy Response
by: Yang, Meng, et al.
Published: (2025) -
Adaptive Multi-resolution Hash-Encoding Framework for INR-based Dental CBCT Reconstruction with Truncated FOV
by: Park, Hyoung Suk, et al.
Published: (2025) -
Neural Image Compression Using Masked Sparse Visual Representation
by: Jiang, Wei, et al.
Published: (2023) -
An Effective Approach to Minimize Error in Midpoint Ellipse Drawing Algorithm
by: Idrisi, M. Javed, et al.
Published: (2021)