PiTe: Pixel-Temporal Alignment for Large Video-Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Yang, Ding, Pengxiang, Huang, Siteng, Zhang, Min, Zhao, Han, Wang, Donglin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference
von: Zhao, Han, et al.
Veröffentlicht: (2024)
von: Zhao, Han, et al.
Veröffentlicht: (2024)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
von: Cui, Can, et al.
Veröffentlicht: (2024)
von: Cui, Can, et al.
Veröffentlicht: (2024)
QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning
von: Tong, Xinyang, et al.
Veröffentlicht: (2024)
von: Tong, Xinyang, et al.
Veröffentlicht: (2024)
Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration
von: Han, Yuhang, et al.
Veröffentlicht: (2024)
von: Han, Yuhang, et al.
Veröffentlicht: (2024)
Prompt-based Distribution Alignment for Unsupervised Domain Adaptation
von: Bai, Shuanghao, et al.
Veröffentlicht: (2023)
von: Bai, Shuanghao, et al.
Veröffentlicht: (2023)
CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
von: Gong, Zhefei, et al.
Veröffentlicht: (2024)
von: Gong, Zhefei, et al.
Veröffentlicht: (2024)
Exploring the Evolution of Physics Cognition in Video Generation: A Survey
von: Lin, Minghui, et al.
Veröffentlicht: (2025)
von: Lin, Minghui, et al.
Veröffentlicht: (2025)
Focus-Consistent Multi-Level Aggregation for Compositional Zero-Shot Learning
von: Dai, Fengyuan, et al.
Veröffentlicht: (2024)
von: Dai, Fengyuan, et al.
Veröffentlicht: (2024)
VGDiffZero: Text-to-image Diffusion Models Can Be Zero-shot Visual Grounders
von: Liu, Xuyang, et al.
Veröffentlicht: (2023)
von: Liu, Xuyang, et al.
Veröffentlicht: (2023)
Rethinking Target Label Conditioning in Adversarial Attacks: A 2D Tensor-Guided Generative Approach
von: Liu, Hangyu, et al.
Veröffentlicht: (2025)
von: Liu, Hangyu, et al.
Veröffentlicht: (2025)
Expressive Forecasting of 3D Whole-body Human Motions
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
von: Cui, Can, et al.
Veröffentlicht: (2025)
von: Cui, Can, et al.
Veröffentlicht: (2025)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
von: Chen, Jiayi, et al.
Veröffentlicht: (2025)
von: Chen, Jiayi, et al.
Veröffentlicht: (2025)
Enhancing Adversarial Transferability via Component-Wise Transformation
von: Liu, Hangyu, et al.
Veröffentlicht: (2025)
von: Liu, Hangyu, et al.
Veröffentlicht: (2025)
Video-Language Alignment via Spatio-Temporal Graph Transformer
von: Zhang, Shi-Xue, et al.
Veröffentlicht: (2024)
von: Zhang, Shi-Xue, et al.
Veröffentlicht: (2024)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
Variation-aware Vision Token Dropping for Faster Large Vision-Language Models
von: Chen, Junjie, et al.
Veröffentlicht: (2025)
von: Chen, Junjie, et al.
Veröffentlicht: (2025)
STaR-KV: Spatio-Temporal Adaptive Re-weighting for KV Cache Compression in GUI Vision-Language Models
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models
von: Huang, Shiqi, et al.
Veröffentlicht: (2026)
von: Huang, Shiqi, et al.
Veröffentlicht: (2026)
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models
von: Song, Wenxuan, et al.
Veröffentlicht: (2026)
von: Song, Wenxuan, et al.
Veröffentlicht: (2026)
Troika: Multi-Path Cross-Modal Traction for Compositional Zero-Shot Learning
von: Huang, Siteng, et al.
Veröffentlicht: (2023)
von: Huang, Siteng, et al.
Veröffentlicht: (2023)
Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
Information-Theoretic Constraints for Continual Vision-Language-Action Alignment
von: Zhao, Libang, et al.
Veröffentlicht: (2026)
von: Zhao, Libang, et al.
Veröffentlicht: (2026)
GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot
von: Song, Wenxuan, et al.
Veröffentlicht: (2024)
von: Song, Wenxuan, et al.
Veröffentlicht: (2024)
Learning Disentangled Identifiers for Action-Customized Text-to-Image Generation
von: Huang, Siteng, et al.
Veröffentlicht: (2023)
von: Huang, Siteng, et al.
Veröffentlicht: (2023)
VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
von: Zhao, Fufangchen, et al.
Veröffentlicht: (2025)
von: Zhao, Fufangchen, et al.
Veröffentlicht: (2025)
PixelLM: Pixel Reasoning with Large Multimodal Model
von: Ren, Zhongwei, et al.
Veröffentlicht: (2023)
von: Ren, Zhongwei, et al.
Veröffentlicht: (2023)
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
von: Fei, Hao, et al.
Veröffentlicht: (2024)
von: Fei, Hao, et al.
Veröffentlicht: (2024)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
von: Wang, Jiankang, et al.
Veröffentlicht: (2025)
von: Wang, Jiankang, et al.
Veröffentlicht: (2025)
PiLoT: Neural Pixel-to-3D Registration for UAV-based Ego and Target Geo-localization
von: Cheng, Xiaoya, et al.
Veröffentlicht: (2026)
von: Cheng, Xiaoya, et al.
Veröffentlicht: (2026)
TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
von: Li, Shicheng, et al.
Veröffentlicht: (2025)
von: Li, Shicheng, et al.
Veröffentlicht: (2025)
VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL
von: Dai, Fengyuan, et al.
Veröffentlicht: (2025)
von: Dai, Fengyuan, et al.
Veröffentlicht: (2025)
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
von: Lu, Yifan, et al.
Veröffentlicht: (2026)
von: Lu, Yifan, et al.
Veröffentlicht: (2026)
On the Consistency of Video Large Language Models in Temporal Comprehension
von: Jung, Minjoon, et al.
Veröffentlicht: (2024)
von: Jung, Minjoon, et al.
Veröffentlicht: (2024)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
von: Fan, Rong, et al.
Veröffentlicht: (2026)
von: Fan, Rong, et al.
Veröffentlicht: (2026)
Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
von: Li, Zeqian, et al.
Veröffentlicht: (2025)
von: Li, Zeqian, et al.
Veröffentlicht: (2025)
Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks
von: Yang, Min, et al.
Veröffentlicht: (2024)
von: Yang, Min, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference
von: Zhao, Han, et al.
Veröffentlicht: (2024) -
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023) -
SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
von: Liu, Yang, et al.
Veröffentlicht: (2025) -
ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
von: Cui, Can, et al.
Veröffentlicht: (2024) -
QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning
von: Tong, Xinyang, et al.
Veröffentlicht: (2024)