Breaking the accuracy-resource dilemma: a lightweight adaptive video inference enhancement
Fuente:
arXiv
Salvato in:
| Autori principali: | Ma, Wei, Chen, Shaowu, Ye, Junjie, Zhang, Peichang, Huang, Lei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PAS: Estimating the target accuracy before domain adaptation
di: Diniz, Raphaella, et al.
Pubblicazione: (2026)
di: Diniz, Raphaella, et al.
Pubblicazione: (2026)
TinyDrop: Tiny Model Guided Token Dropping for Vision Transformers
di: Wang, Guoxin, et al.
Pubblicazione: (2025)
di: Wang, Guoxin, et al.
Pubblicazione: (2025)
ARIW-Framework: Adaptive Robust Iterative Watermarking Framework
di: Wu, Shaowu, et al.
Pubblicazione: (2025)
di: Wu, Shaowu, et al.
Pubblicazione: (2025)
Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
di: Chen, Lei, et al.
Pubblicazione: (2025)
di: Chen, Lei, et al.
Pubblicazione: (2025)
DiTraj: training-free trajectory control for video diffusion transformer
di: Lei, Cheng, et al.
Pubblicazione: (2025)
di: Lei, Cheng, et al.
Pubblicazione: (2025)
video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM
di: Sun, Guangzhi, et al.
Pubblicazione: (2025)
di: Sun, Guangzhi, et al.
Pubblicazione: (2025)
Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta Compression
di: Wang, Xiaohui, et al.
Pubblicazione: (2025)
di: Wang, Xiaohui, et al.
Pubblicazione: (2025)
A lightweight detector for real-time detection of remote sensing images
di: Wang, Qianyi, et al.
Pubblicazione: (2025)
di: Wang, Qianyi, et al.
Pubblicazione: (2025)
Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models
di: Zhong, Yi, et al.
Pubblicazione: (2026)
di: Zhong, Yi, et al.
Pubblicazione: (2026)
Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence
di: Granite Vision Team, et al.
Pubblicazione: (2025)
di: Granite Vision Team, et al.
Pubblicazione: (2025)
MCVI-SANet: A lightweight semi-supervised model for LAI and SPAD estimation of winter wheat under vegetation index saturation
di: Zhang, Zhiheng, et al.
Pubblicazione: (2025)
di: Zhang, Zhiheng, et al.
Pubblicazione: (2025)
MemeBLIP2: A novel lightweight multimodal system to detect harmful memes
di: Liu, Jiaqi, et al.
Pubblicazione: (2025)
di: Liu, Jiaqi, et al.
Pubblicazione: (2025)
AC-MAMBASEG: An adaptive convolution and Mamba-based architecture for enhanced skin lesion segmentation
di: Nguyen, Viet-Thanh, et al.
Pubblicazione: (2024)
di: Nguyen, Viet-Thanh, et al.
Pubblicazione: (2024)
Can video generation replace cinematographers? Research on the cinematic language of generated video
di: Li, Xiaozhe, et al.
Pubblicazione: (2024)
di: Li, Xiaozhe, et al.
Pubblicazione: (2024)
Geospatial foundation models for image analysis: evaluating and enhancing NASA-IBM Prithvi's domain adaptability
di: Hsu, Chia-Yu, et al.
Pubblicazione: (2024)
di: Hsu, Chia-Yu, et al.
Pubblicazione: (2024)
Warfare:Breaking the Watermark Protection of AI-Generated Content
di: Li, Guanlin, et al.
Pubblicazione: (2023)
di: Li, Guanlin, et al.
Pubblicazione: (2023)
Uncovering Intrinsic Capabilities: A Paradigm for Data Curation in Vision-Language Models
di: Li, Junjie, et al.
Pubblicazione: (2025)
di: Li, Junjie, et al.
Pubblicazione: (2025)
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
di: Zheng, Mingzhe, et al.
Pubblicazione: (2024)
di: Zheng, Mingzhe, et al.
Pubblicazione: (2024)
A motion-based compression algorithm for resource-constrained video camera traps
di: Ratnayake, Malika Nisal, et al.
Pubblicazione: (2024)
di: Ratnayake, Malika Nisal, et al.
Pubblicazione: (2024)
Phantom: Subject-consistent video generation via cross-modal alignment
di: Liu, Lijie, et al.
Pubblicazione: (2025)
di: Liu, Lijie, et al.
Pubblicazione: (2025)
Flow caching for autoregressive video generation
di: Ma, Yuexiao, et al.
Pubblicazione: (2026)
di: Ma, Yuexiao, et al.
Pubblicazione: (2026)
Dermatologist-like explainable AI enhances melanoma diagnosis accuracy: eye-tracking study
di: Chanda, Tirtha, et al.
Pubblicazione: (2024)
di: Chanda, Tirtha, et al.
Pubblicazione: (2024)
Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling
di: Fu, Xiaolong, et al.
Pubblicazione: (2025)
di: Fu, Xiaolong, et al.
Pubblicazione: (2025)
Breaking the Data Barrier -- Building GUI Agents Through Task Generalization
di: Zhang, Junlei, et al.
Pubblicazione: (2025)
di: Zhang, Junlei, et al.
Pubblicazione: (2025)
Breast tumor classification based on self-supervised contrastive learning from ultrasound videos
di: Tang, Yunxin, et al.
Pubblicazione: (2024)
di: Tang, Yunxin, et al.
Pubblicazione: (2024)
Breaking the Barrier: Selective Uncertainty-based Active Learning for Medical Image Segmentation
di: Ma, Siteng, et al.
Pubblicazione: (2024)
di: Ma, Siteng, et al.
Pubblicazione: (2024)
FFA Sora, video generation as fundus fluorescein angiography simulator
di: Wu, Xinyuan, et al.
Pubblicazione: (2024)
di: Wu, Xinyuan, et al.
Pubblicazione: (2024)
Unsupervised Domain Adaptation via Similarity-based Prototypes for Cross-Modality Segmentation
di: Ye, Ziyu, et al.
Pubblicazione: (2025)
di: Ye, Ziyu, et al.
Pubblicazione: (2025)
CLIPPan: Adapting CLIP as A Supervisor for Unsupervised Pansharpening
di: Jian, Lihua, et al.
Pubblicazione: (2025)
di: Jian, Lihua, et al.
Pubblicazione: (2025)
Fine-Grained Urban Flow Inference with Multi-scale Representation Learning
di: Yuan, Shilu, et al.
Pubblicazione: (2024)
di: Yuan, Shilu, et al.
Pubblicazione: (2024)
WebForge: Breaking the Realism-Reproducibility-Scalability Trilemma in Browser Agent Benchmark
di: Yuan, Peng, et al.
Pubblicazione: (2026)
di: Yuan, Peng, et al.
Pubblicazione: (2026)
Geometric Knowledge-Guided Localized Global Distribution Alignment for Federated Learning
di: Ma, Yanbiao, et al.
Pubblicazione: (2025)
di: Ma, Yanbiao, et al.
Pubblicazione: (2025)
Visual moral inference and communication
di: Zhu, Warren, et al.
Pubblicazione: (2025)
di: Zhu, Warren, et al.
Pubblicazione: (2025)
Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents
di: Ma, Tianyi, et al.
Pubblicazione: (2025)
di: Ma, Tianyi, et al.
Pubblicazione: (2025)
Breaking Self-Attention Failure: Rethinking Query Initialization for Infrared Small Target Detection
di: Liu, Yuteng, et al.
Pubblicazione: (2026)
di: Liu, Yuteng, et al.
Pubblicazione: (2026)
MaxGlaViT: A novel lightweight vision transformer-based approach for early diagnosis of glaucoma stages from fundus images
di: Yurdakul, Mustafa, et al.
Pubblicazione: (2025)
di: Yurdakul, Mustafa, et al.
Pubblicazione: (2025)
One-shot synthesis of rare gastrointestinal lesions improves diagnostic accuracy and clinical training
di: Yu, Jia, et al.
Pubblicazione: (2025)
di: Yu, Jia, et al.
Pubblicazione: (2025)
Learning from Pattern Completion: Self-supervised Controllable Generation
di: Chen, Zhiqiang, et al.
Pubblicazione: (2024)
di: Chen, Zhiqiang, et al.
Pubblicazione: (2024)
ID-Guard: A Universal Framework for Combating Facial Manipulation via Breaking Identification
di: Qu, Zuomin, et al.
Pubblicazione: (2024)
di: Qu, Zuomin, et al.
Pubblicazione: (2024)
IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression Segmentation
di: Chen, Qi, et al.
Pubblicazione: (2025)
di: Chen, Qi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
PAS: Estimating the target accuracy before domain adaptation
di: Diniz, Raphaella, et al.
Pubblicazione: (2026) -
TinyDrop: Tiny Model Guided Token Dropping for Vision Transformers
di: Wang, Guoxin, et al.
Pubblicazione: (2025) -
ARIW-Framework: Adaptive Robust Iterative Watermarking Framework
di: Wu, Shaowu, et al.
Pubblicazione: (2025) -
Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
di: Chen, Lei, et al.
Pubblicazione: (2025) -
DiTraj: training-free trajectory control for video diffusion transformer
di: Lei, Cheng, et al.
Pubblicazione: (2025)