Atom: Efficient On-Device Video-Language Pipelines Through Modular Reuse
Fuente:
arXiv
Saved in:
| Main Authors: | Panchal, Kunjal, Mitra, Saayan, Sarkhel, Somdeb, Wang, Haoliang, Dasgupta, Ishita, Wu, Gang, Guan, Hui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SKALD: Learning-Based Shot Assembly for Coherent Multi-Shot Video Creation
by: Lu, Chen Yi, et al.
Published: (2025)
by: Lu, Chen Yi, et al.
Published: (2025)
A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations
by: Deshpande, Vijeta, et al.
Published: (2025)
by: Deshpande, Vijeta, et al.
Published: (2025)
Flow: Per-Instance Personalized Federated Learning Through Dynamic Routing
by: Panchal, Kunjal, et al.
Published: (2022)
by: Panchal, Kunjal, et al.
Published: (2022)
Thinking Forward: Memory-Efficient Federated Finetuning of Language Models
by: Panchal, Kunjal, et al.
Published: (2024)
by: Panchal, Kunjal, et al.
Published: (2024)
The Cost of Avoiding Backpropagation
by: Panchal, Kunjal, et al.
Published: (2025)
by: Panchal, Kunjal, et al.
Published: (2025)
E2LVLM:Evidence-Enhanced Large Vision-Language Model for Multimodal Out-of-Context Misinformation Detection
by: Wu, Junjie, et al.
Published: (2025)
by: Wu, Junjie, et al.
Published: (2025)
ELLMPEG: An Edge-based Agentic LLM Video Processing Tool
by: Azimi, Zoha, et al.
Published: (2026)
by: Azimi, Zoha, et al.
Published: (2026)
FlexCache: Flexible Approximate Cache System for Video Diffusion
by: Sun, Desen, et al.
Published: (2024)
by: Sun, Desen, et al.
Published: (2024)
Machine Learning-Based Prediction of Quality Shifts on Video Streaming Over 5G
by: Mustafa, Raza Ul, et al.
Published: (2025)
by: Mustafa, Raza Ul, et al.
Published: (2025)
Group Benefits Instances Selection for Data Purification
by: Cai, Zhenhuang, et al.
Published: (2024)
by: Cai, Zhenhuang, et al.
Published: (2024)
A Clustering-Based Method for Automatic Educational Video Recommendation Using Deep Face-Features of Lecturers
by: Mendes, Paulo R. C., et al.
Published: (2020)
by: Mendes, Paulo R. C., et al.
Published: (2020)
Efficient Distributed Training through Gradient Compression with Sparsification and Quantization Techniques
by: Singh, Shruti, et al.
Published: (2024)
by: Singh, Shruti, et al.
Published: (2024)
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
by: Qian, Wenhao, et al.
Published: (2025)
by: Qian, Wenhao, et al.
Published: (2025)
QoS-QoE Translation with Large Language Model
by: Yu, Yingjie, et al.
Published: (2026)
by: Yu, Yingjie, et al.
Published: (2026)
Predicting Outcomes in Video Games with Long Short Term Memory Networks
by: Chulajata, Kittimate, et al.
Published: (2024)
by: Chulajata, Kittimate, et al.
Published: (2024)
Multimodal Representation Learning and Fusion
by: Jin, Qihang, et al.
Published: (2025)
by: Jin, Qihang, et al.
Published: (2025)
Enhancing Cross-Prompt Transferability in Vision-Language Models through Contextual Injection of Target Tokens
by: Yang, Xikang, et al.
Published: (2024)
by: Yang, Xikang, et al.
Published: (2024)
OOD-GraphLLM: Graph Large Language Model for Out-of-Distribution Generalized Drug Synergy Prediction
by: Wang, Xin, et al.
Published: (2026)
by: Wang, Xin, et al.
Published: (2026)
DeepTextMark: A Deep Learning-Driven Text Watermarking Approach for Identifying Large Language Model Generated Text
by: Munyer, Travis, et al.
Published: (2023)
by: Munyer, Travis, et al.
Published: (2023)
Coherent Audio-Visual Editing via Conditional Audio Generation Following Video Edits
by: Ishii, Masato, et al.
Published: (2025)
by: Ishii, Masato, et al.
Published: (2025)
Hybrid Feedback-Guided Optimal Learning for Wireless Interactive Panoramic Scene Delivery
by: Wu, Xiaoyi, et al.
Published: (2026)
by: Wu, Xiaoyi, et al.
Published: (2026)
Beyond Interpretability: Exploring the Comprehensibility of Adaptive Video Streaming through Large Language Models
by: Jia, Lianchen, et al.
Published: (2025)
by: Jia, Lianchen, et al.
Published: (2025)
Demonstration of MaskSearch: Efficiently Querying Image Masks for Machine Learning Workflows
by: Wei, Lindsey Linxi, et al.
Published: (2024)
by: Wei, Lindsey Linxi, et al.
Published: (2024)
HyperFusion: Hierarchical Multimodal Ensemble Learning for Social Media Popularity Prediction
by: Ye, Liliang, et al.
Published: (2025)
by: Ye, Liliang, et al.
Published: (2025)
L3GS: Layered 3D Gaussian Splats for Efficient 3D Scene Delivery
by: Tsai, Yi-Zhen, et al.
Published: (2025)
by: Tsai, Yi-Zhen, et al.
Published: (2025)
Video DataFlywheel: Resolving the Impossible Data Trinity in Video-Language Understanding
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
Group Cognition Learning: Making Everything Better Through Governed Two-Stage Agents Collaboration
by: Meng, Chunlei, et al.
Published: (2026)
by: Meng, Chunlei, et al.
Published: (2026)
NeuSaver: Neural Adaptive Power Consumption Optimization for Mobile Video Streaming
by: Park, Kyoungjun, et al.
Published: (2021)
by: Park, Kyoungjun, et al.
Published: (2021)
Automatic Camera Trajectory Control with Enhanced Immersion for Virtual Cinematography
by: Wu, Xinyi, et al.
Published: (2023)
by: Wu, Xinyi, et al.
Published: (2023)
LinVT: Empower Your Image-level Large Language Model to Understand Videos
by: Gao, Lishuai, et al.
Published: (2024)
by: Gao, Lishuai, et al.
Published: (2024)
Harmful Visual Content Manipulation Matters in Misinformation Detection Under Multimedia Scenarios
by: Wang, Bing, et al.
Published: (2026)
by: Wang, Bing, et al.
Published: (2026)
AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction
by: Xing, Zhen, et al.
Published: (2024)
by: Xing, Zhen, et al.
Published: (2024)
Music Genre Classification: Ensemble Learning with Subcomponents-level Attention
by: Liu, Yichen, et al.
Published: (2024)
by: Liu, Yichen, et al.
Published: (2024)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
by: Li, Quanhao, et al.
Published: (2025)
by: Li, Quanhao, et al.
Published: (2025)
MCAD: Multimodal Context-Aware Audio Description Generation For Soccer
by: Chaudhary, Lipisha, et al.
Published: (2025)
by: Chaudhary, Lipisha, et al.
Published: (2025)
Deconfounded Reasoning for Multimodal Fake News Detection via Causal Intervention
by: Liu, Moyang, et al.
Published: (2025)
by: Liu, Moyang, et al.
Published: (2025)
Balanced Multimodal Learning: An Unidirectional Dynamic Interaction Perspective
by: Wang, Shijie, et al.
Published: (2025)
by: Wang, Shijie, et al.
Published: (2025)
Multimodal Methods for Analyzing Learning and Training Environments: A Systematic Literature Review
by: Cohn, Clayton, et al.
Published: (2024)
by: Cohn, Clayton, et al.
Published: (2024)
OpenAVS: Training-Free Open-Vocabulary Audio Visual Segmentation with Foundational Models
by: Chen, Shengkai, et al.
Published: (2025)
by: Chen, Shengkai, et al.
Published: (2025)
Cross-Space Synergy: A Unified Framework for Multimodal Emotion Recognition in Conversation
by: Lyu, Xiaosen, et al.
Published: (2025)
by: Lyu, Xiaosen, et al.
Published: (2025)
Similar Items
-
SKALD: Learning-Based Shot Assembly for Coherent Multi-Shot Video Creation
by: Lu, Chen Yi, et al.
Published: (2025) -
A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations
by: Deshpande, Vijeta, et al.
Published: (2025) -
Flow: Per-Instance Personalized Federated Learning Through Dynamic Routing
by: Panchal, Kunjal, et al.
Published: (2022) -
Thinking Forward: Memory-Efficient Federated Finetuning of Language Models
by: Panchal, Kunjal, et al.
Published: (2024) -
The Cost of Avoiding Backpropagation
by: Panchal, Kunjal, et al.
Published: (2025)