VA-$π$: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liao, Xinyao, He, Qiyuan, Xu, Kai, Qu, Xiaoye, Li, Yicong, Wei, Wei, Yao, Angela |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
von: Chen, Kewei, et al.
Veröffentlicht: (2025)
von: Chen, Kewei, et al.
Veröffentlicht: (2025)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
von: Chen, Kewei, et al.
Veröffentlicht: (2026)
von: Chen, Kewei, et al.
Veröffentlicht: (2026)
Physics-informed Variational Autoencoders for Improved Robustness to Environmental Factors of Variation
von: Thoreau, Romain, et al.
Veröffentlicht: (2022)
von: Thoreau, Romain, et al.
Veröffentlicht: (2022)
Tricks and Plug-ins for Gradient Boosting in Image Classification
von: Fang, Biyi, et al.
Veröffentlicht: (2025)
von: Fang, Biyi, et al.
Veröffentlicht: (2025)
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
von: Tu, Songjun, et al.
Veröffentlicht: (2025)
von: Tu, Songjun, et al.
Veröffentlicht: (2025)
CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method
von: Koh, Hyunseo, et al.
Veröffentlicht: (2026)
von: Koh, Hyunseo, et al.
Veröffentlicht: (2026)
Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think
von: Tian, Jie, et al.
Veröffentlicht: (2025)
von: Tian, Jie, et al.
Veröffentlicht: (2025)
Unpacking Hateful Memes: Presupposed Context and False Claims
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
Training a Student Expert via Semi-Supervised Foundation Model Distillation
von: Taghavi, Pardis, et al.
Veröffentlicht: (2026)
von: Taghavi, Pardis, et al.
Veröffentlicht: (2026)
The Power of Next-Frame Prediction for Learning Physical Laws
von: Winterbottom, Thomas, et al.
Veröffentlicht: (2024)
von: Winterbottom, Thomas, et al.
Veröffentlicht: (2024)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
von: Qesaraku, Bjorna, et al.
Veröffentlicht: (2025)
von: Qesaraku, Bjorna, et al.
Veröffentlicht: (2025)
Do Generative Metrics Predict YOLO Performance? An Evaluation Across Models, Augmentation Ratios, and Dataset Complexity
von: Marian, Vasile, et al.
Veröffentlicht: (2026)
von: Marian, Vasile, et al.
Veröffentlicht: (2026)
Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling
von: Jung, Seoik, et al.
Veröffentlicht: (2025)
von: Jung, Seoik, et al.
Veröffentlicht: (2025)
Training for X-Ray Vision: Amodal Segmentation, Amodal Content Completion, and View-Invariant Object Representation from Multi-Camera Video
von: Moore, Alexander, et al.
Veröffentlicht: (2025)
von: Moore, Alexander, et al.
Veröffentlicht: (2025)
Cooperative Perception: A Resource-Efficient Framework for Multi-Drone 3D Scene Reconstruction Using Federated Diffusion and NeRF
von: Pourmandi, Massoud
Veröffentlicht: (2025)
von: Pourmandi, Massoud
Veröffentlicht: (2025)
Visible and Hyperspectral Imaging for Quality Assessment of Milk: Property Characterisation and Identification
von: Martinelli, Massimo, et al.
Veröffentlicht: (2026)
von: Martinelli, Massimo, et al.
Veröffentlicht: (2026)
Balanced conic rectified flow
von: Kim, Shin Seong, et al.
Veröffentlicht: (2025)
von: Kim, Shin Seong, et al.
Veröffentlicht: (2025)
SemanticFeels: Semantic Labeling during In-Hand Manipulation
von: Khalil, Anas Al Shikh, et al.
Veröffentlicht: (2026)
von: Khalil, Anas Al Shikh, et al.
Veröffentlicht: (2026)
IDOL: Instant Photorealistic 3D Human Creation from a Single Image
von: Zhuang, Yiyu, et al.
Veröffentlicht: (2024)
von: Zhuang, Yiyu, et al.
Veröffentlicht: (2024)
TACIT: Transformation-Aware Capturing of Implicit Thought
von: Nobrega, Daniel
Veröffentlicht: (2026)
von: Nobrega, Daniel
Veröffentlicht: (2026)
Salient Concept-Aware Generative Data Augmentation
von: Zhao, Tianchen, et al.
Veröffentlicht: (2025)
von: Zhao, Tianchen, et al.
Veröffentlicht: (2025)
Spectral Integrated Gradients for Coarse-to-Fine Feature Attribution
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
Manifold-Aligned Guided Integrated Gradients for Reliable Feature Attribution
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
Method of UAV Inspection of Photovoltaic Modules Using Thermal and RGB Data Fusion
von: Lysyi, Andrii, et al.
Veröffentlicht: (2025)
von: Lysyi, Andrii, et al.
Veröffentlicht: (2025)
ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
von: Syed, Shahram Najam, et al.
Veröffentlicht: (2025)
von: Syed, Shahram Najam, et al.
Veröffentlicht: (2025)
Supervised Learning Has a Necessary Geometric Blind Spot: Theory, Consequences, and Minimal Repair
von: Rajput, Vishal
Veröffentlicht: (2026)
von: Rajput, Vishal
Veröffentlicht: (2026)
AUTHENTICATION: Identifying Rare Failure Modes in Autonomous Vehicle Perception Systems using Adversarially Guided Diffusion Models
von: Zarei, Mohammad, et al.
Veröffentlicht: (2025)
von: Zarei, Mohammad, et al.
Veröffentlicht: (2025)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
von: Sáez, Arnau Igualde, et al.
Veröffentlicht: (2025)
von: Sáez, Arnau Igualde, et al.
Veröffentlicht: (2025)
A Landmark-Aware Visual Navigation Dataset
von: Johnson, Faith, et al.
Veröffentlicht: (2024)
von: Johnson, Faith, et al.
Veröffentlicht: (2024)
Rethinking Visual Intelligence: Insights from Video Pretraining
von: Acuaviva, Pablo, et al.
Veröffentlicht: (2025)
von: Acuaviva, Pablo, et al.
Veröffentlicht: (2025)
Multimodal Generative AI for Story Point Estimation in Software Development
von: Islam, Mohammad Rubyet, et al.
Veröffentlicht: (2025)
von: Islam, Mohammad Rubyet, et al.
Veröffentlicht: (2025)
Surg$Σ$: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence
von: Zeng, Zhitao, et al.
Veröffentlicht: (2026)
von: Zeng, Zhitao, et al.
Veröffentlicht: (2026)
Contrastive Consolidation of Top-Down Modulations Achieves Sparsely Supervised Continual Learning
von: Tran, Viet Anh Khoa, et al.
Veröffentlicht: (2025)
von: Tran, Viet Anh Khoa, et al.
Veröffentlicht: (2025)
Biomedical Visual Instruction Tuning with Clinician Preference Alignment
von: Cui, Hejie, et al.
Veröffentlicht: (2024)
von: Cui, Hejie, et al.
Veröffentlicht: (2024)
Sat-JEPA-Diff: Bridging Self-Supervised Learning and Generative Diffusion for Remote Sensing
von: Komurcu, Kursat, et al.
Veröffentlicht: (2026)
von: Komurcu, Kursat, et al.
Veröffentlicht: (2026)
ForAug: Recombining Foregrounds and Backgrounds to Improve Vision Transformer Training with Bias Mitigation
von: Nauen, Tobias Christian, et al.
Veröffentlicht: (2025)
von: Nauen, Tobias Christian, et al.
Veröffentlicht: (2025)
A Survey on Vision-Language-Action Models for Embodied AI
von: Ma, Yueen, et al.
Veröffentlicht: (2024)
von: Ma, Yueen, et al.
Veröffentlicht: (2024)
Measuring What Matters: Scenario-Driven Evaluation for Trajectory Predictors in Autonomous Driving
von: Da, Longchao, et al.
Veröffentlicht: (2025)
von: Da, Longchao, et al.
Veröffentlicht: (2025)
Learning Association via Track-Detection Matching for Multi-Object Tracking
von: Adžemović, Momir
Veröffentlicht: (2025)
von: Adžemović, Momir
Veröffentlicht: (2025)
Convolutional Model Trees
von: Armstrong, William Ward, et al.
Veröffentlicht: (2025)
von: Armstrong, William Ward, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
von: Chen, Kewei, et al.
Veröffentlicht: (2025) -
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
von: Chen, Kewei, et al.
Veröffentlicht: (2026) -
Physics-informed Variational Autoencoders for Improved Robustness to Environmental Factors of Variation
von: Thoreau, Romain, et al.
Veröffentlicht: (2022) -
Tricks and Plug-ins for Gradient Boosting in Image Classification
von: Fang, Biyi, et al.
Veröffentlicht: (2025) -
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
von: Tu, Songjun, et al.
Veröffentlicht: (2025)