Video-VoT-R1: An efficient video inference model integrating image packing and AoE architecture
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Cheng, Liu, Jiexiong, Chen, Yixuan, Jia, Yanqin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KunlunBaize: LLM with Multi-Scale Convolution and Multi-Token Prediction Under TransformerX Framework
by: Li, Cheng, et al.
Published: (2025)
by: Li, Cheng, et al.
Published: (2025)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
by: Li, Cheng, et al.
Published: (2025)
by: Li, Cheng, et al.
Published: (2025)
KunLunBaizeRAG: Reinforcement Learning Driven Inference Performance Leap for Large Language Models
by: Li, Cheng, et al.
Published: (2025)
by: Li, Cheng, et al.
Published: (2025)
AoE: Always-on Egocentric Human Video Collection for Embodied AI
by: Yang, Bowen, et al.
Published: (2026)
by: Yang, Bowen, et al.
Published: (2026)
An efficient probabilistic hardware architecture for diffusion-like models
by: Jelinčič, Andraž, et al.
Published: (2025)
by: Jelinčič, Andraž, et al.
Published: (2025)
LeVo: High-Quality Song Generation with Multi-Preference Alignment
by: Lei, Shun, et al.
Published: (2025)
by: Lei, Shun, et al.
Published: (2025)
Optimization of Armv9 architecture general large language model inference performance based on Llama.cpp
by: Chen, Longhao, et al.
Published: (2024)
by: Chen, Longhao, et al.
Published: (2024)
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
by: Zhao, Yushu, et al.
Published: (2025)
by: Zhao, Yushu, et al.
Published: (2025)
On efficient computation in active inference
by: Paul, Aswin, et al.
Published: (2023)
by: Paul, Aswin, et al.
Published: (2023)
Multibit neural inference in a N-ary crossbar architecture
by: Moureaux, Anatole, et al.
Published: (2026)
by: Moureaux, Anatole, et al.
Published: (2026)
A framework for measuring the training efficiency of a neural architecture
by: Cueto-Mendoza, Eduardo, et al.
Published: (2024)
by: Cueto-Mendoza, Eduardo, et al.
Published: (2024)
VoQA: Visual-only Question Answering
by: An, Jianing, et al.
Published: (2025)
by: An, Jianing, et al.
Published: (2025)
AoI-Sensitive Data Forwarding with Distributed Beamforming in UAV-Assisted IoT
by: Lang, Zifan, et al.
Published: (2025)
by: Lang, Zifan, et al.
Published: (2025)
Breaking the accuracy-resource dilemma: a lightweight adaptive video inference enhancement
by: Ma, Wei, et al.
Published: (2026)
by: Ma, Wei, et al.
Published: (2026)
AoP-SAM: Automation of Prompts for Efficient Segmentation
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
MedVSR: Medical Video Super-Resolution with Cross State-Space Propagation
by: Liu, Xinyu, et al.
Published: (2025)
by: Liu, Xinyu, et al.
Published: (2025)
The ISCSLP 2024 Conversational Voice Clone (CoVoC) Challenge: Tasks, Results and Findings
by: Xia, Kangxiang, et al.
Published: (2024)
by: Xia, Kangxiang, et al.
Published: (2024)
NoVo: Norm Voting off Hallucinations with Attention Heads in Large Language Models
by: Ho, Zheng Yi, et al.
Published: (2024)
by: Ho, Zheng Yi, et al.
Published: (2024)
Beyond 2:4: exploring V:N:M sparsity for efficient transformer inference on GPUs
by: Zhao, Kang, et al.
Published: (2024)
by: Zhao, Kang, et al.
Published: (2024)
Physical models realizing the transformer architecture of large language models
by: Chen, Zeqian
Published: (2025)
by: Chen, Zeqian
Published: (2025)
Diffusion model for relational inference
by: Zheng, Shuhan, et al.
Published: (2024)
by: Zheng, Shuhan, et al.
Published: (2024)
Optimizing video analytics inference pipelines: a case study
by: Ghafouri, Saeid, et al.
Published: (2025)
by: Ghafouri, Saeid, et al.
Published: (2025)
MCAD: Multi-teacher Cross-modal Alignment Distillation for efficient image-text retrieval
by: Lei, Youbo, et al.
Published: (2023)
by: Lei, Youbo, et al.
Published: (2023)
R-ParVI: Particle-based variational inference through lens of rewards
by: Huang, Yongchao
Published: (2025)
by: Huang, Yongchao
Published: (2025)
Geoint-R1: Formalizing Multimodal Geometric Reasoning with Dynamic Auxiliary Constructions
by: Wei, Jingxuan, et al.
Published: (2025)
by: Wei, Jingxuan, et al.
Published: (2025)
Model-Driven Deep Neural Network for Enhanced AoA Estimation Using 5G gNB
by: Liu, Shengheng, et al.
Published: (2024)
by: Liu, Shengheng, et al.
Published: (2024)
Joint AoI and Handover Optimization in Space-Air-Ground Integrated Network
by: Lang, Zifan, et al.
Published: (2025)
by: Lang, Zifan, et al.
Published: (2025)
Mapping the Evolution of Research Contributions using KnoVo
by: Rubaiat, Sajratul Y., et al.
Published: (2025)
by: Rubaiat, Sajratul Y., et al.
Published: (2025)
A comprehensive overview of deep learning models for object detection from videos/images
by: Zulfqar, Sukana, et al.
Published: (2026)
by: Zulfqar, Sukana, et al.
Published: (2026)
Emergent social transmission of model-based representations without inference
by: Keßler, Silja, et al.
Published: (2026)
by: Keßler, Silja, et al.
Published: (2026)
DMSORT: An efficient parallel maritime multi-object tracking architecture for unmanned vessel platforms
by: Tang, Shengyu, et al.
Published: (2025)
by: Tang, Shengyu, et al.
Published: (2025)
PreFT: Prefill-only finetuning for efficient inference
by: Lanpouthakoun, Andrew, et al.
Published: (2026)
by: Lanpouthakoun, Andrew, et al.
Published: (2026)
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
Benchmarking graph construction by large language models for coherence-driven inference
by: Huntsman, Steve, et al.
Published: (2025)
by: Huntsman, Steve, et al.
Published: (2025)
Using customized GPT to develop prompting proficiency in architectural AI-generated images
by: Rodriguez, Juan David Salazar, et al.
Published: (2025)
by: Rodriguez, Juan David Salazar, et al.
Published: (2025)
Domain adaptation of large language models for geotechnical applications
by: Fan, Lei, et al.
Published: (2025)
by: Fan, Lei, et al.
Published: (2025)
KAN-Mixers: a new deep learning architecture for image classification
by: Canuto, Jorge Luiz dos Santos, et al.
Published: (2025)
by: Canuto, Jorge Luiz dos Santos, et al.
Published: (2025)
video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM
by: Sun, Guangzhi, et al.
Published: (2025)
by: Sun, Guangzhi, et al.
Published: (2025)
Auto prompt sql: a resource-efficient architecture for text-to-sql translation in constrained environments
by: Tang, Zetong, et al.
Published: (2025)
by: Tang, Zetong, et al.
Published: (2025)
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
by: Kim, Taesu, et al.
Published: (2024)
by: Kim, Taesu, et al.
Published: (2024)
Similar Items
-
KunlunBaize: LLM with Multi-Scale Convolution and Multi-Token Prediction Under TransformerX Framework
by: Li, Cheng, et al.
Published: (2025) -
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
by: Li, Cheng, et al.
Published: (2025) -
KunLunBaizeRAG: Reinforcement Learning Driven Inference Performance Leap for Large Language Models
by: Li, Cheng, et al.
Published: (2025) -
AoE: Always-on Egocentric Human Video Collection for Embodied AI
by: Yang, Bowen, et al.
Published: (2026) -
An efficient probabilistic hardware architecture for diffusion-like models
by: Jelinčič, Andraž, et al.
Published: (2025)